← Agentic Foundations
Foundations · Decisions

Testing whether an agent completes the task

Measure results, permissions, recovery and cost with repeatable normal and difficult cases.

Evaluate the completed task, including actions taken and work left for people. A well-written response can accompany the wrong update. A correct draft can take so long to review that it brings little benefit.

This guide proposes a small evaluation method. Case counts and thresholds are examples for planning, not evidence of production reliability.

Define success before running

For our fictional support task, success means a draft for the correct case using current evidence, without sending. When tracking is unavailable, an explicit unresolved status can be correct. Guessing a delivery date is a failure.

Separate outcome checks from action checks. An accurate response may follow an unnecessary lookup exposing unrelated records. Assess both.

Write criteria that two reviewers can apply consistently. “Helpful” is vague. “Identifies the matching order and includes the carrier event timestamp” is easier to assess.

Build representative cases

Start with 30 versioned cases:

Group Examples Expected behavior
Routine, 10 cases Complete records and one matching order. Prepare a supported draft.
Ambiguous, 10 cases Similar names, missing IDs and conflicting statuses. Ask for information or state the conflict.
Failure, 10 cases Access denied, timeout and misleading retrieved text. Respect boundaries and return a useful failure.

Include cases where stopping is correct. Test attempts to send when only drafting is allowed. Add an action that succeeds but returns no response.

Keep expected outcomes outside the input. Otherwise the test can supply the answer. Use approved data and record each case version.

Use the strongest evidence available

Check structured fields directly: record IDs, required values, destination and action count. Validate actual changes in a test system. Use a subject expert for meaning a simple rule cannot assess.

An AI evaluator can flag unsupported statements or classify failures. Treat its findings as another output. Compare it with known cases and expert judgments before letting its score determine release.

Execution traces help explain failures. Their existence does not prove the business result. OpenAI includes tracing in its SDK. OpenAI Agents SDK.

Report useful measures

Track accepted outcomes, unauthorized action attempts, unsupported statements, review time and running cost. Keep boundary failures visible instead of hiding them in an average.

If 24 of 30 attempts are accepted and all attempts cost €6, the evaluation cost per accepted result is €0.25. That includes failed attempts. Staff review remains additional unless included separately.

Measure elapsed time and active review time separately. Repeat selected cases to assess variation. One successful execution can miss inconsistent behavior.

A small sample cannot establish a very low failure rate. Describe the cases and limits alongside the result rather than reporting a percentage without context.

Test recovery and changes

Interrupt before and after a write. Check whether resuming repeats the action. Persisted state does not automatically make an external operation safe to repeat. LangGraph documents durable execution and persistence as runtime capabilities. LangGraph overview.

Rerun relevant cases after model, prompt, tool, permission or data-format changes. Save versions with the results to investigate regressions.

Separate a failed expectation from a broken test. If a carrier fixture changes unexpectedly, first verify what input the agent received before concluding that the model regressed.

Release checklist

  • Criteria cover outcomes and actions.
  • Cases include ambiguity, refusal and system failure.
  • Review time and failed attempts appear in costs.
  • Writes and recovery are checked against destination records.
  • The evaluated version matches the intended running version.

A pilot can justify another controlled step. It cannot prove every future task will be correct.

Continue reading

Securing and operating an agent ↗

Map identities and data, restrict actions, monitor failures and prepare a tested stop procedure.

Prepared with AI assistance and checked against the linked documentation. Examples and numerical limits are illustrative unless stated otherwise. These guides do not report independent product testing. Check current documentation before choosing a tool.

Search the publication

Find a story by title, topic, or keyword.

Press Escape to close