Library · topic

Evaluation

Agents fail differently from chat models: trajectories, not single outputs, and nondeterminism that hides in pass@1. These notes cover eval pitfalls, repeated-trial metrics, retrieval-vs-generation attribution, and the latency/cost dimensions that decide whether an agent is usable, not just correct.

Ask a question about evaluation