Offline and online agent evaluation
Added 2026-07-14 · Last verified 2026-07-14
Offline evals provide repeatable pre-release regression tests over representative tasks, adversarial cases, tool failures, and multi-step trajectories; online evals measure real completion, escalation, latency, cost, and safety signals under production traffic. Maintain frozen golden sets plus newly mined failures, run repeated trials for nondeterministic agents, and score both final outcomes and critical intermediate constraints. LLM judges are scalable but can be biased by style, verbosity, ordering, or shared model errors, so calibrate them against blinded human labels and deterministic checks and track judge-version changes.
Related notes
Source: The Agent Loop — agent-loop.xyz. Curated in the open knowledge base (CC BY 4.0).