← Agentic Foundations
Foundations · Basics

Using an agent and checking its work

Give a clear task, understand its permissions and verify the result before you rely on it.

When you use an agent, you delegate part of a task. You need to understand the expected result, which actions it can take and how you will know what it completed. You do not need to learn an SDK to ask those questions.

This guide offers a suggested working method. The example uses a fictional support case and assumes your organization has approved the application and its data access.

Give it a task you can check

A useful request names the task, records, permitted actions and output:

Check delivery status for case SC-1042 using the order and tracking records. Prepare a reply with the latest confirmed status and list any missing information. Do not send the reply or change the order.

The case identifier narrows scope. The output distinguishes facts from missing information. Preparation and sending are separate actions.

Avoid adding passwords, unnecessary customer details or sensitive documents. Use the approved access method. A long prompt cannot fix an application that exposes more data than the task needs.

Know the actions before starting

Ask the owner whether your agent can only read and draft or can also change records and send messages. Check whether it acts with your permissions or another identity. A chat interface does not establish read-only access.

Some developer tools provide approval controls. The Claude Agent SDK exposes permission modes, rules and runtime decisions for tool requests. Whether your application asks you for approval depends on its implementation. Claude Agent SDK permissions.

If the interface cannot explain what an action affects, ask the owner before assigning a consequential task.

Review the evidence and the output

Question Useful evidence
Is this the correct case? Case and order identifiers match.
How current is the status? The tracking event has a date and time.
Did it reconcile conflicting records? It explains the disagreement.
Did it send anything? The action record shows drafting without sending.
What remains unresolved? Missing tracking is listed rather than guessed.

Read cited records when the result affects a customer or business decision. An agent can attach the right source and misunderstand it. Check the statement against the evidence.

“Completed” should describe the agreed task. A useful response might say: “Draft prepared; delivery date remains unconfirmed because the carrier lookup failed.”

Handle errors without repeating the damage

If a task times out after a write, verify the destination before running it again. A missing response does not prove that nothing changed. Repeating a send operation could produce another message.

Report the task identifier, expected behavior and observed result through the approved support channel. You rarely need an entire customer conversation to explain that a lookup used the wrong order.

Check a proposed correction against the same evidence. A confident second answer still needs verification.

Make the workflow accessible

You should be able to inspect sources, approve actions and report problems using your normal access needs. Important status information should appear as readable text, not only as a color or temporary notification.

Ask for an alternative when a step depends on an inaccessible document, visual-only chart or pointer-only interaction. Include that requirement in pilot feedback so the owner can change the process.

Before you rely on the result

  • Confirm task and record identifiers.
  • Check source dates and missing information.
  • Review consequential drafts before sending.
  • Verify changes in the destination system.
  • Report mistakes with a task identifier and specific example.

For a coding example of the same principle, GitHub advises reviewing and testing generated work because it can contain errors and vulnerabilities. GitHub documentation.

Continue reading

Who owns the decisions around an agent? ↗

Assign responsibility for outcomes, actions, access, evaluation and incidents before work begins.

Prepared with AI assistance and checked against the linked documentation. Examples and numerical limits are illustrative unless stated otherwise. These guides do not report independent product testing. Check current documentation before choosing a tool.

Search the publication

Find a story by title, topic, or keyword.

Press Escape to close