Read freely, write carefully: controls for AI agents that act
An agent that reads can mislead you. An agent that writes or sends can do something you cannot take back. The controls should reflect the difference.
Most early uses of AI are reading tasks. Summarise this. Search that. Draft a reply. When the model is wrong, the harm is a wrong answer, and there is a chance that a person catches it before it goes anywhere.
Agents change that. An agent calls tools. It updates records, raises tickets and sends messages. A mistake is no longer a bad paragraph. It is an action.
The OWASP Top 10 for LLM Applications calls this risk Excessive Agency. It describes a vulnerability that lets a system take damaging actions in response to unexpected, ambiguous or manipulated output from a model. The usual root causes are excessive functionality, excessive permissions and excessive autonomy. In the current 2026 edition it is LLM03, up from LLM06 in 2025. The controls below are one practical response.
Permission to build is not permission to deploy
Several separate decisions hide inside the question “can we use an agent for this?”
Connecting a data source is one decision. Granting the agent access to a system is another. Sharing the tool with colleagues is a third. Letting it act without a person in the loop is a fourth. Each carries different risks, and each may need a different owner to agree to it.
A prototype that works well on a test inbox is evidence that the idea can work. It is not approval to run it on a shared mailbox. Record each decision on its own, so that nobody mistakes one for another.
Read freely, write carefully
Classify every tool the agent can call by its effect.
- Read: no side effects. Search, fetch, summarise.
- Write: changes something inside your own systems. Update a record, create a draft.
- Send: reaches outside. Email, publish, notify, pay.
- Irreversible: anything you cannot cleanly undo.
Reads can be broadly allowed, but not blindly. A read can put sensitive data in front of the model, and a later send can carry it out. Writes need limits. Sends and irreversible actions need the strongest controls. That is the rule in the title.
Keep the list of callable actions explicit. If an action is not on the allow-list, the agent cannot call it. Give each tool the narrowest function and the narrowest permissions it needs. OWASP’s mitigations for Excessive Agency make the same points. They add one that is easy to miss: enforce authorisation in code, outside the model, rather than letting the model decide whether an action is allowed. The 2026 entry also describes graduated enforcement. Low-consequence or easily reversible actions can be approved automatically, and high-consequence or irreversible ones go to a person.
Four modes: allow, draft, approve, block
Give each type of action a mode.
- Allow: the action runs directly. Suitable for reads and low-risk writes.
- Draft: the agent prepares the action, and nothing happens until a person takes it forward themselves.
- Approve: the agent proposes a specific action, and a named person approves exactly that action before it runs.
- Block: the action never runs. This is the default for anything not on the allow-list.
Approval has to be of the real thing. Show the approver the actual recipients, the actual content and the actual values, not the agent’s summary of them. Approving a summary means approving the agent’s description of what it will do, which may not match what it does.
Where a draft is enough, prefer it to approval. A draft gives a tired reviewer nothing to wave through.
Re-validate at execution time
Time passes between proposal and execution. In that gap, things change. A distribution list is edited. A record is updated by someone else. A permission is withdrawn. The kill switch is pulled.
So check again at the moment of execution. The action must still be allowed and within limits. The approval must not have expired, and the kill switch must be off. The payload must be exactly what was approved. Comparing a hash of the approved payload with a hash of what is about to run checks that cheaply. If any check fails, the action does not run.
Make retries safe
Agents retry. Networks fail. A timeout does not mean the action did not happen. Without care, a retry sends the same message twice.
Give each action a key derived from the tool and its arguments. Normalise the arguments first, so that trivial differences such as spacing or the order of recipients do not count. Record the key before the action runs. A second request with the same key, within a set window, is refused as a duplicate. That turns “has this already gone?” from a guess into a lookup.
Failures are the hard case. If a call timed out, the message may or may not have gone. Refusing every retry turns a passing network fault into a permanent refusal. Allowing retries risks a duplicate. Choose deliberately, and write the choice down.
Keep an audit trail you can trust
Record what was proposed, and by which agent. Record who approved it and when, what actually ran, and what came back. Include the inputs that influenced the action: if a document contained instructions that steered the agent, you will want to find it later. Leave secrets out.
Make the log append-only, and chain each entry to the hash of the one before. An edit, an insertion, a reordering or a deletion from the middle then breaks the chain, and a check can find it. Two things do not break it: cutting off the newest entries, and rewriting the whole file with new hashes. So keep a copy of the latest hash somewhere the agent cannot write to. The log is then tamper-evident. It is not tamper-proof.
Have a kill switch, and test it
There should be one control that stops every action at once, including actions a person has already approved. The component that executes actions should check it before every action, rather than asking the agent to respect it. Keep approved actions in the queue while it is on, so they can be reviewed rather than lost. A switch in code stops only what passes through it, so in a real incident revoke the agent’s credentials as well.
A kill switch that has never been used is a hope, not a control. Test it on a schedule. Pair it with rate limits and spending budgets. OWASP lists rate limits and circuit breakers among the measures that will not prevent Excessive Agency but can limit the damage.
Before any of this
Ask whether you need an agent at all. Many processes improve more from cleaner data, a better template or ordinary automation than from a model that decides what to do next. Fix the data first. Add an agent when the remaining problem really needs judgement, and give it only the reach the task requires.
One of the ten principles in the AI Playbook for the UK Government is that you have “meaningful human control at the right stages”. That is easier to write than to build. Controls like these are part of what it looks like in code.
Sources
- OWASP GenAI Security Project, LLM03:2026 Excessive Agency, OWASP Top 10 for LLM Applications 2026 (source text on GitHub). See “Complete mediation” and “Rate limiting” under prevention and mitigation.
- OWASP GenAI Security Project, LLM06:2025 Excessive Agency, OWASP Top 10 for LLM Applications 2025 (source text on GitHub).
- AI Playbook for the UK Government, GOV.UK. See principle 4, “You have meaningful human control at the right stages”.
Views are my own. Where I refer to public guidance, this is my reading of it: unofficial; not government guidance.