没跑过 = 不出厂。no eval, no ship.
-
Updated
Sep 15, 2026
没跑过 = 不出厂。no eval, no ship.
Chat agentico multi-agente em Streamlit: um CEO delega, especialistas executam e um juiz valida as respostas.
Production RAG system — hybrid search, Cohere reranking, RAGAS regression gate, Langfuse tracing, and confidence guardrails. Eval-gated deployment on AWS Lambda.
Governed, auditable AI review platform: the review and approval work a team does by hand, as a webhook-triggered LangGraph agent with human-in-the-loop approval, Langfuse tracing, an eval-gated CI pipeline, and a tamper-evident hash-chained audit trail. Provider-agnostic (mock/Anthropic/OpenAI), FastAPI, Bicep IaC for Azure Container Apps.
An LLM gateway that learns from your evals to route smarter, cheaper.
Public SOP for enterprise AI adoption: write the boundary before you automate. Diagnose, access, engineer, deliver, compound.
A reference architecture for evidence-bounded generative media pipelines — grounded, eval-gated, with a fail-closed-to-publish model. Curated extract; ships no third-party data.
Field notes from building AI systems in production since 2019. LLM serving ops: cost tracking, eval-driven deployment, capacity planning, observability, incident playbooks.
To associate your repository with the eval-driven topic, visit your repo's landing page and select "manage topics."