AI Agent Tool Permissions: A Side-Effect and Approval Envelope

AI agent tool permissions should describe the effect an operation may produce, the resources it may touch and the conditions that permit execution. Human approval belongs inside that boundary. An approval envelope records the exact proposed action, its permitted lifetime and the circumstances that require the system to stop or ask again.
Start with an operation that matters. Consider a hypothetical renewal assistant that reads an assigned customer account, prepares an internal note and proposes an email to the existing account contact. The assistant can help with all three activities, but the authority needed for each differs. Editing a private draft does not justify sending it, and permission to send one message does not justify exporting the account history.
The design below is an engineering worksheet for that example, not a documented client deployment or a universal security standard. Its purpose is to make the boundary inspectable before a production credential reaches the agent. The same review can expose where an apparently reversible operation produces an external consequence that a database rollback cannot erase.
Classify the effect that another system can observe
Begin with the downstream effect rather than the tool's friendly name. Account management might include retrieving a record, changing an internal note, adding a recipient and deleting the account. Those operations should not inherit the same permission because they happen to share a connector.
A useful first register separates the following activities. The approval choices are proposed defaults for the renewal example. The product owner must adapt them to the actual data sensitivity and business authority.
| Operation | Observable effect | Proposed permission boundary | Approval and recovery consequence |
|---|---|---|---|
| Read assigned account | Account data enters the agent context | Current tenant and assigned account only | No per-read approval inside approved scope. Restrict onward disclosure |
| Prepare private draft | New content exists in an internal workspace | Draft store and permitted fields | Usually automatic. Retain source references and discard capability |
| Edit internal review note | A shared record changes | Named fields and expected record version | Permit bounded edits only if notifications and dependent jobs are understood |
| Send renewal email | A recipient receives content | Exact recipient, body and attachment set | Require explicit approval. Delivery cannot reliably be undone |
| Export account documents | Copies leave their original access boundary | Named documents and approved destination | Separate disclosure approval. Deletion of the export does not recall copies |
| Grant access or delete account | Authority or availability changes | Excluded from this assistant's tools | Route to a separately authorized workflow. No conversational override |
Read-only describes the operation at the source. It says nothing about the destination of the retrieved information.
A read that places confidential text into an unauthorized external service still creates a disclosure problem. Include allowed destinations in the register even when the source credential cannot write.
Reversibility also needs a precise definition. An internal note may be restored, yet its earlier value might already have triggered a notification or another worker. List those observers before labeling the operation reversible. If their consequences cannot be recalled, evaluate the complete effect as an externally visible change.
OWASP's Excessive Agency guidance recommends minimizing tool functionality and downstream permissions, using human approval for high-impact actions and enforcing authorization outside the language model. The register turns that guidance into decisions about this assistant's specific operations. It does not establish that the implementation already meets those decisions.
Build an approval envelope around one proposal
For the email operation, create a structured proposal before asking for approval. It should be possible to inspect the proposal without rereading the entire agent conversation. A stable action reference connects the visible review to the object that the execution service will receive.
| Envelope field | Example constraint | What invalidates the proposal |
|---|---|---|
| Principal and tenant | Renewal workflow acting for an authorized account owner | Different identity or tenant |
| Operation and resource | Send one message concerning the assigned account | Another operation or account |
| Recipient | Existing verified account contact | Address change, added recipient or forwarding destination |
| Payload | Reviewed subject, body and attachment identifiers | Material content or attachment change |
| Preconditions | Account version observed when the proposal was prepared | Relevant account state has changed |
| Approver | Named person with authority for this communication | Authority revoked or approval supplied by an ineligible person |
| Validity | Explicit expiry and one permitted execution | Expiry, prior consumption or cancellation |
| Policy identity | Version of the applicable permission rules | A policy change that requires reassessment |
| Recovery rule | Hold uncertain delivery for reconciliation | A new send proposed as an unexamined retry |
These fields form an application contract, not an instruction that the model is trusted to enforce. Normalize the fields whose representation can vary, then bind the approval to that normalized proposal. Define the normalization rules explicitly: silently dropping a recipient or attachment while computing the reference would bind approval to the wrong effect.
A digest can help detect a changed proposal. It does not establish who approved it or protect an editable approval store by itself.
Store the approval with authenticated identity, controlled write access and the policy decision that allowed that person to approve. Keep the human-readable proposal available so an operator can interpret the digest's meaning.
Decide what counts as a material edit. A change to harmless display spacing need not create a new business decision if the canonical content is unchanged. Changing a price statement, recipient or attachment does. The rule belongs in the sending workflow. An agent should not get to declare its own changes immaterial.
Keep business approval separate from connector authority
Approval answers whether an otherwise permitted action should happen in this instance. Downstream authorization answers whether the executing identity may perform that operation against that resource now. Both decisions must succeed before the governed effect occurs.
Give the renewal assistant separate read and send surfaces where the integration permits it. The reader cannot gain mailbox-wide write capability merely because an email draft was approved. A sending service should accept the approved proposal and expose the minimum operation needed to execute it. Its credential remains constrained by the destination system's permissions.
Some providers offer scopes that are broader than the intended business operation. Document that mismatch.
An application gate can reject disallowed requests through its own path, but a stolen broad credential may bypass that gate. Credential storage and downstream access restrictions therefore remain part of the review. Do not describe a narrow tool schema as eliminating a broader credential's exposure.
Verify tenant and resource ownership at the enforcing component. An account identifier supplied by the model is an input to validate, not evidence that the account belongs to the requester. Prefer resource resolution from an authenticated tenant context over accepting a free-form tenant value from the proposed call.
If authorization cannot be established, keep the write pending or deny it under the agreed policy. A dependency outage should not become implicit permission. The user interface can explain that the proposal remains prepared while execution is held. This preserves useful work without turning a service failure into an authority expansion.
Show the approver the consequence, not the whole transcript
The approval screen should display what will change and who will observe it.
For the renewal email, show the resolved recipient, the final message and the complete attachment set. Put source facts beside consequential statements so the reviewer can distinguish account data from an agent's inference.
Anthropic states: Agents can then pause for human feedback at checkpoints or when encountering blockers.
This sentence appears in Building effective agents, published by Anthropic on December 19, 2024 and credited to Erik S. and Barry Zhang. It supports a checkpoint in the workflow. It does not claim that an approval button provides an authorization system.
Make the review object stable while the approval is being considered. If account data changes, tell the reviewer which facts are stale and require the relevant reassessment. An approval screen that continues showing the old message while the executor sends a revised version defeats the purpose of the checkpoint.
For a changed proposal, display the difference from the previously reviewed version. The reviewer should not need to spot an added attachment inside a long block of text. A small recipient change can alter disclosure scope more than a large stylistic rewrite, so order the visible differences by consequence.
Avoid treating every low-impact read as an approval event. Repeated prompts consume attention that the consequential decision needs. Automatic reads can remain within a previously authorized narrow scope, while external communication gets an inspectable checkpoint. The operation register should explain that distinction so fewer prompts represent a deliberate policy decision.
Recheck the envelope immediately before dispatch
An approved proposal can become stale while waiting in a queue. The contact address may change, the account owner may lose authority or the customer may already have received a message from another workflow. Execution must consult the current state relevant to those conditions.
Record the preconditions that matter to the effect. A version check can reject a stale record update. A send operation may instead need a fresh contact resolution and a check for an already completed business action. Do not assume that one record version captures changes across every system involved.
Use a dispatch gate that checks current permission, proposal integrity and applicable approval. Consume a single-use approval through an atomic local transition so two workers cannot both take it as unused. That transition prevents duplicate dispatch through the controlled path. It does not, by itself, guarantee one downstream effect across a network failure.
Revocation has a timing boundary. If permission is removed before the gate allows dispatch, the action should stop. A request already accepted by a remote provider may finish after local revocation. Describe that in-flight limitation and the supported downstream cancellation behavior rather than promising immediate recall.
Choose expiry to match how long the proposal's facts remain useful. A short lifetime creates more review work. A long lifetime increases the chance that circumstances change. The right value requires workflow evidence. In a pilot, record queue delay and stale-proposal rejection before widening the lifetime, and identify who may authorize that change.
Distinguish replay permission from an uncertain result
After dispatch, a timeout can leave delivery unknown. The approval was consumed, but the sending provider may already have accepted the message. Returning the action to an unused approval state would make a second send appear legitimate without resolving the first.
Maintain separate execution states: prepared, approved, dispatched, confirmed, denied and uncertain. A confirmed result needs downstream evidence appropriate to the operation, such as a created message reference. An uncertain result needs an investigation owner and a permitted query path. It is neither proof of success nor proof that nothing happened.
AWS describes caller-provided request identifiers in Making retries safe with idempotent APIs. Its examples distinguish repeating the same intent from issuing a genuinely new request. Apply that distinction only where the actual provider documents compatible idempotency behavior, including the identifier's scope and retention window.
Preserve the business action reference during a supported retry. Generating a fresh identifier after a timeout can turn an attempted replay into a new operation. Conversely, reusing an identifier for a materially changed message can confuse two different intentions. The proposal binding and the provider's idempotency contract must agree on what remains the same.
When the provider lacks a usable deduplication or lookup mechanism, retain the uncertainty. A human may need to reconcile the message history before authorizing another attempt. The decision should cite the evidence inspected and the remaining ambiguity. A stalled action is operationally inconvenient, but inventing evidence of non-delivery hides the risk of another effect.
Treat compensation as a separately permitted action
Compensation repairs an outcome. It rarely erases every consequence.
Restoring an internal note can correct the current record while leaving a notification in a person's inbox. Sending a correction can reduce confusion while leaving the original email in circulation.
For each write operation, describe the repair available and what it cannot restore. The renewal assistant might restore a previous draft automatically within its workspace, but it should not send a correction to a customer under permission that covered only the original message. That communication needs its own proposal and authority.
Keep the repair path narrow. An operator investigating an uncertain email should be able to inspect the relevant delivery evidence without receiving unrestricted send access. If reconciliation reveals the need for a new message, move to the normal approval path with a new action reference.
The repair record should explain the original effect, the intended correction and the expected remaining exposure. Identify the person who owns communication with affected users. A technical rollback result is not evidence that those users never saw the earlier state.
Some actions have no satisfactory compensation within the product. Exported documents can be copied onward. Deletion may remove information required for recovery. Exclude those operations from an early assistant if the business cannot define an acceptable response. Expanding autonomy should be a separate design decision supported by new controls and evidence.
Bound the aggregate effect of a composed workflow
Several individually permitted actions can exceed the intended task. Reading one assigned account and sending one approved email is different from searching all accounts, collecting documents and contacting every discovered person. Evaluate the workflow's accumulated effect as well as each call.
Specify aggregate allowances where repetition changes the risk: permitted recipients, total documents disclosed and executions per approved task. An operation count is useful only when its unit is clear. One tool call might send a batch of messages, so a limit expressed solely in calls can conceal the actual number of recipients.
Approval of a batch should cover a fixed membership or a precisely defined rule whose consequences the reviewer can inspect. In the renewal example, a later discovered contact should not silently join an approved recipient set. A material expansion creates a new proposal even if the message template stays identical.
Delegation needs the same boundary. A second agent may prepare content using information already permitted for the task, but it should not inherit a broader credential or manufacture approval on the first agent's behalf. Treat its returned proposal as untrusted input to the original execution gate.
Stop conditions also belong to the aggregate policy. If the workflow reaches its disclosure allowance or receives contradictory resource state, pause the affected operation and explain the unresolved condition. Do not split the task into new runs solely to evade the allowance. The business owner must explicitly define when a genuinely new task begins.
Test invalid approvals and forbidden effects
Successful tool execution proves little about the boundary. Acceptance tests should attempt the changes the envelope forbids and inspect whether a downstream request was prevented. Use a controlled environment and synthetic data for cases that could otherwise expose information or contact real recipients.
| Test condition | Expected decision in this design | Evidence to inspect |
|---|---|---|
| Recipient changes after approval | Hold for a new proposal and approval | Changed field, rejected action reference and no governed dispatch |
| Approval expires in the queue | Do not execute the send | Expiry decision and retained pending proposal |
| Request targets another tenant | Deny at the enforcing boundary | Resolved tenant, denial and downstream request absence |
| Two workers consume one approval | Only one local dispatch reservation succeeds | Atomic transition records and attempted worker references |
| Provider response is lost after dispatch | Mark uncertain and reconcile | Attempt reference, available provider evidence and no blind new send |
| Access is revoked before dispatch | Block the affected operation | Current authorization result and blocked proposal |
Do not stop at checking that an error appeared in the agent transcript.
Inspect the component responsible for the side effect and the receiving system where feasible. A model saying it refused is insufficient if another execution route can still call the provider.
Include instruction-bearing source content in the tests. For example, an account note may tell the assistant to add a new recipient or disregard a previous limit. The note remains task data. It should not change the authenticated policy or become approval merely because it appears in a retrieved record.
Record the build and policy version covered by each result. Recheck affected cases when adding an operation, broadening credentials or changing proposal normalization. Keep unrun cases visible as unverified. A demonstration with an old sending configuration does not establish the boundary of the new deployment.
Account for control work before widening autonomy
Human review and reconciliation are recurring operating work. Model them separately from model-token consumption. A design that adds a checkpoint without an owner, queue capacity or escalation path can leave a useful assistant waiting indefinitely for attention.
Use explicit assumptions. Suppose a pilot prepares 600 proposals each month, 20% need human review and each review takes 90 seconds. That produces 120 reviews and 180 minutes, or three hours, of review time. These figures are illustrative inputs, not measured productivity or a recommended approval rate.
Now suppose 12 proposals require reconciliation, averaging 15 minutes each. That adds another 180 minutes. The example therefore allocates six hours across review and reconciliation before counting policy maintenance, access reviews or incident work. If review duration doubles, review alone becomes six hours and the combined allocation becomes nine. Replace the inputs with observed pilot records.
The control worksheet should identify what the delivery team must implement: operation separation, proposal binding, authorization checks and result reconciliation. Pharos Production's AI development services describe discovery, workflow integration and production delivery for AI systems. For a write-capable assistant, make the permission envelope part of that scope and agree its acceptance evidence before expanding the tool surface. The service description establishes a relevant offering. This worksheet does not claim a delivered security outcome.
Assign ownership for policy changes and stale approvals, then identify who handles uncertain results. The NIST AI Risk Management Framework provides voluntary risk-management context. Using it as a reference does not certify the assistant or establish compliance with a particular law.
Approval evidence also has a data boundary. Preserve the record needed to reconstruct the decision, but avoid copying whole account histories into a log by default. Store sensitive reviewed content under appropriate access controls and expose redacted references to routine operators. Decide its retention period with the responsible data owner. An uneditable receipt is useful only if authorized investigators can retrieve the relevant evidence without creating another uncontrolled disclosure.
Begin with the first externally visible operation your agent will perform. Write its proposal fields, identify the enforcing gate and rehearse a changed recipient plus a lost response. Keep the capability closed until those outcomes are explainable from the records. That gives the next permission expansion a concrete starting point: one known effect, bounded authority and an observed recovery path.
Evidence notes
Sources checked on : the linked Anthropic article, OWASP Excessive Agency guidance, AWS Builders’ Library paper, NIST AI Risk Management Framework and AI service description. The renewal assistant, approval envelope and acceptance cases are original illustrative design recommendations. Workload figures are fictional assumptions, not observed client results, market benchmarks or guaranteed security outcomes.
Comments
Post a Comment