← Agentic Foundations
Foundations · Decisions

Choosing a task and running a useful pilot

Define the business result, limit the first task and compare the agent with your current process.

Start your pilot with a task and an acceptance decision. “Try agents in customer service” leaves too many possible goals. “Prepare a delivery-status draft for a verified case” gives you inputs, boundaries and a result to measure.

This method is editorial guidance. Its numbers are illustrative planning choices, not a standard or a performance promise.

Decide whether the task needs an agent

Write down the steps people perform today. Identify where the sequence changes and why. If the same lookup and message work every time, a fixed workflow may suffice. If staff must select sources and respond to incomplete results, model-directed decisions might help.

Microsoft distinguishes open-ended tasks from processes needing explicit execution order. It recommends using a function when a function can handle the job. Microsoft Agent Framework overview.

An agent becomes a candidate when you can explain the flexibility it needs and how you will check its choices. A task with no reliable way to assess the result makes a poor first pilot.

Write a one-page task agreement

Item Fictional support pilot decision
Result A draft with confirmed status or an explicit unresolved question.
Inputs Verified case, linked order and permitted carrier record.
Allowed work Read those records and create a private draft.
Excluded actions Sending replies, issuing refunds and editing orders.
Failure route Return the case to the queue with the failed step.
Owner A named support manager responsible for accepting the result.

Add expected service hours and a way to stop processing. Confirm who handles corrections. These decisions should exist before the pilot receives live work.

Measure the current process first

Observe representative cases. Record active staff time, elapsed time, correctness and repair work. Separate waiting time from work time. A slow carrier system can make a task take an hour even when staff spend only five minutes on it.

Include incomplete records, ambiguous identifiers and ordinary failures. A collection of clean examples can make a weak system appear reliable.

Limit personal data to what the assessment needs. Fabricated cases with realistic structure may suffice for an early prototype. Confirm approved data handling before using real cases.

Compare the same outcome

Suppose you evaluate 50 cases: 30 routine, 10 incomplete and 10 conflicting. Apply the same acceptance criteria to the manual and agent-assisted processes.

A draft must refer to the correct order, cite the actual status and avoid an unapproved send. Substantial rewriting does not count as clean success. Include review time in process cost.

Reduced drafting time is useful only if quality remains acceptable. Include service errors and user experience. Ask staff whether evidence is easy to inspect and whether failure messages help them finish the work.

Evaluate accessibility with the people who use the process. A faster draft offers little value if its approval screen cannot be operated with their access needs.

Decide before expanding

Agree which results mean continue, change or stop. You might require every test to respect the no-send boundary and return ambiguous identities for review. A small sample with zero failures does not prove the system will never fail.

Treat broader scope as another decision. Sending replies changes the consequence of a mistake. Another language changes your cases. A different model changes previously assessed behavior.

Document who authorizes the next stage and what evidence they require. Keep unresolved failures visible in that decision.

Pilot checklist

  • One task, one accountable owner and explicit excluded actions.
  • A baseline using the same acceptance criteria.
  • Routine, ambiguous and failure cases.
  • Review time, repair time and running costs included.
  • A stop procedure, decision date and criteria for expansion.

Continue with the evaluation guide to turn these decisions into reproducible checks.

Continue reading

Who owns the decisions around an agent? ↗

Assign responsibility for outcomes, actions, access, evaluation and incidents before work begins.

Prepared with AI assistance and checked against the linked documentation. Examples and numerical limits are illustrative unless stated otherwise. These guides do not report independent product testing. Check current documentation before choosing a tool.

Search the publication

Find a story by title, topic, or keyword.

Press Escape to close