Skip to content
View garceslabs's full-sized avatar

Block or report garceslabs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
garceslabs/README.md

Garces Labs

Production AI Systems • Decision Intelligence • Multi-Agent Orchestration • AI Reliability

Building production-grade AI systems that combine LLM reasoning with deterministic software engineering. Focused on evaluation, orchestration, reliability, and decision support for real-world AI applications.

Former Google GenAI Manager leading multimodal evaluation, AI quality systems, and production AI launches at global scale.


Focus Areas

  • Multi-Agent Orchestration
  • AI Decision Intelligence
  • LLM Evaluation & Reliability
  • Production AI Infrastructure
  • Human-in-the-Loop AI
  • Deterministic AI Workflows
  • AI Systems for Decision Support

Current Interests

  • Agentic AI Systems
  • AI Evaluation & Benchmarking
  • Frontier Model Reliability
  • Decision Intelligence Platforms
  • Algorithmic Trading & Investment Systems
  • Multi-Agent Architectures
  • AI Developer Tooling

Flagship Projects

🚀 Application Intelligence Platform

A multi-stage AI system that evaluates job opportunities, tailors truthful CVs, performs hiring-manager reviews, benchmarks AI providers, validates outputs, and generates deterministic application artifacts.

Highlights

  • Provider abstraction (Codex CLI, Claude CLI, OpenAI API)
  • Deterministic orchestration
  • Evidence-linked contracts
  • Benchmark framework
  • Human-in-the-loop decision support

📈 North Star Capital

A multi-agent investment decision platform that simulates an institutional investment committee through specialized analyst roles, deterministic orchestration, evidence-based reasoning, and explainable portfolio recommendations.

Highlights

  • 17-role investment committee
  • Two-round deliberation framework
  • Evidence validation
  • Risk-aware decision synthesis
  • Explainable investment recommendations

🧪 LLM Evals Platform

Evaluation framework for hallucination detection, factuality, uncertainty scoring, multimodal consistency validation, and long-horizon task reliability.


🤖 AI Coordination System

Operational framework for orchestrating AI workflows, release gating, evaluation pipelines, escalation management, and production AI quality systems.


📚 Research Orchestrator Agent

Evidence-driven research assistant focused on retrieval, contradiction detection, source ranking, and citation-aware synthesis.


Engineering Philosophy

Reliable AI systems are built by combining strong model capabilities with deterministic orchestration, rigorous evaluation, and human oversight.

The objective is not simply to generate outputs, but to build AI systems that are observable, reproducible, and trustworthy in production.


Tech Stack

PythonFastAPILLMsAgentic SystemsEvaluation FrameworksStructured OutputsPytestDockerCI/CDOpenAIAnthropic


Current Goal

Building AI systems that bridge frontier models and production software engineering through evaluation, orchestration, benchmarking, and deterministic workflows.

Expanding these principles into decision intelligence across domains including AI evaluation, hiring systems, research automation, and algorithmic investment workflows.

Connect

  • LinkedIn: link
  • Technical writing: coming soon

Pinned Loading

  1. llm-evals-platform llm-evals-platform Public

    Evaluation infrastructure for hallucination, jailbreak, and factuality testing of LLM systems.

    Python

  2. ai-coordination-system ai-coordination-system Public

    Production-grade coordination layer for multi-agent AI systems, task routing, escalation management, and human-in-the-loop reliability.

    Python

  3. research-orchestrator-agent research-orchestrator-agent Public

    Production-style research agent focused on grounded reasoning, source attribution, factuality, and verification workflows.

    Python