An agent can say that a refund is complete while the customer's account shows no refund. For an agent that can use tools, a convincing reply is only one piece of evidence. A production evaluation has to check the actions the agent took and the world those actions left behind.
Teams are already using production conversations and human labels to build agent evaluations, as Dropbox's engineering account illustrates; the durable problem is deciding whether an agent actually completed an authorized task. How do we turn that question into a pipeline that can test a new agent version before release and keep learning afterward?
The compact answer is a loop: observed request → reviewed case → isolated trial → trace and final state → graders → version comparison → release decision → production feedback. A case describes the starting world and acceptable outcomes. A trial runs one agent attempt against that world. Graders inspect independent evidence, and the comparison asks whether the new version improves the tasks that matter without breaking important slices. The loop is reusable; the meaning of success changes with the agent's job.
The first useful rule is to give a completed action its own proof. If a support agent says “refunded,” a ledger record or service commit is the proof, subject to the service's consistency contract. If the record is absent, the polished sentence cannot rescue the task. That single separation prevents an answer-only judge from turning a failed action into a pass.
The second rule is to test the task from a known starting point. The same request can require a refund when the customer verifies identity, a clarifying question when they cannot, or a refusal when policy forbids it. An evaluation case must carry those starting facts and accepted branches. Otherwise a grader may punish a careful agent for asking or reward a reckless one for acting.
Before trusting a score, ask what its grader can observe. If it sees only messages, the score measures answer quality. If it sees tool calls but no final state, it can verify attempted actions without proving durable effects. Reserve “task success” for a test whose evidence covers the task's actual outcome. This distinction is useful even before the full pipeline exists.
Table of contents




