An agent with tools can access data and perform actions. Assess the application, identities, tools and hosting together. The model provider is one part of that system.
This guide supplies implementation and operating questions. It is not a legal assessment or evidence of compliance. Confirm data requirements with your organization’s appropriate owners.
Map identities and access
Record the identity used for each operation. A user can sign into the interface while a backend uses a service credential. The backend must decide which records that user may request.
Limit access to the task. A support-drafting agent does not need organization-wide write permission. Separate read, draft and send capabilities so you can withdraw one independently.
Keep credentials in an approved secret store. Avoid tokens in prompts, descriptions, logs or artifacts. Record who rotates them and how running jobs receive changes.
Follow the data
List what travels to models, tool servers, storage and monitoring. Include attachments, conversation history, tool results and traces. A useful trace can contain sensitive content.
Record purpose, access, retention and deletion for each destination. Confirm provider terms and the actual deployment configuration. Services from the same vendor can have different processing arrangements.
A lookup needing an order ID may not need an account export. Redaction should preserve necessary evidence while removing unnecessary personal information.
Treat retrieved content as untrusted
An email or page can ask the agent to disclose data or run another tool. Receiving that text does not authorize the action. Separate source content from application instructions and validate actions at execution time.
Test a carrier note asking the agent to send customer information elsewhere. The expected result is to treat it as content and stay within the task permissions.
MCP’s security guidance covers risks including confused-deputy behavior, token handling and requests to unintended network destinations. A protocol does not remove those risks. MCP security best practices.
Restrict the execution environment
For file or shell tools, define directories, commands and network destinations. Runtime permissions and operating-system isolation address different parts of the problem. Assess both when the agent can execute code.
For business APIs, validate identifiers and limits in backend code. Approval for consequential actions should show exactly what will change.
Claude’s SDK exposes tool permission controls. Your application still needs a suitable policy. Claude Agent SDK permissions.
Do not give a runtime access to production credentials simply because it needs a realistic test. Create a separate test identity with controlled data and permissions.
Monitor the task, not only uptime
A healthy server can produce incorrect drafts. Track successful tasks, failed steps, denied actions, duplicate operations and usage costs. Connect events with a task ID without collecting unnecessary content.
Alert on meaningful conditions such as repeated failures, unexpected writes or a budget limit. Assign investigators and a route for user-reported mistakes.
Record model, prompt, tool and policy versions. Confirm which configuration caused a problem before applying a correction.
Prepare an incident procedure
Write a runbook: pause new tasks, disable affected actions, inspect destination records, preserve appropriate evidence and notify responsible owners. Identify tasks needing correction.
Stopping execution does not reverse completed actions. A sent message or transaction needs a separate recovery process.
Test whether queued and running tasks stop, whether credentials remain usable elsewhere and whether restarting repeats unfinished writes. Decide who can authorize resuming.
Operating checklist
- Document identities and permissions per tool.
- Confirm data destinations and retention.
- Prevent retrieved text from granting new authority.
- Constrain execution and backend actions.
- Assign owners for investigation, pausing and recovery.
Reassess after access, tool or task changes.
Testing whether an agent completes the task ↗
Measure results, permissions, recovery and cost with repeatable normal and difficult cases.
Prepared with AI assistance and checked against the linked documentation. Examples and numerical limits are illustrative unless stated otherwise. These guides do not report independent product testing. Check current documentation before choosing a tool.