Notes on agentic AI, evals, infrastructure, and shipping LLM systems in production.
Retrieval and generation get all the attention. Evaluation is where you find out if any of it actually works — here's how to build that layer properly.