When an AI agent does something it shouldn't: approval gates and blast-radius limits

The support agent cleared 200 tickets overnight, and everyone was glad until one turned out to be a refund no customer had asked for. Nobody could say what triggered it. By the next day the agent was switched off, and a project that looked finished on Friday was back to arguing whether an agent could be trusted with real data.
That argument usually starts in the wrong place. An agent that only reads and drafts is low-stakes even when wrong; one that can refund, send, or delete carries a different risk, because a single wrong action lands on real customers and records. AI agent guardrails keep that action small and recoverable rather than a reason to shut it down.
You can give an agent real authority without betting the business on its judgment. Grant it the minimum tools and access its job needs, require a human to approve anything irreversible, and cap what any single action can touch. Log every action with its reason, so a mistake stays visible and undoable.
The controls below are how experienced engineers ship an acting agent to production: separating autonomy from authority, limiting its reach, gating the one-way doors, and bounding the blast radius.
What AI agent guardrails actually protect against
AI agent guardrails exist to contain the worst single action an agent can take unsupervised. For an agent that only reads, that worst case is a weak suggestion. For one that can refund, send, or delete, it is an action on real data that a customer or a colleague feels immediately, and a run of good behaviour beforehand does nothing to soften it. The risk scales with what the agent can touch rather than with how often it is right.
This failure has a name. The OWASP project lists Excessive Agency among the top risks for LLM applications, and traces its root cause to "excessive functionality, excessive permissions, or excessive autonomy" — an agent given more reach than its task needs. The controls in this article map to those three causes directly: limit what the agent can do, limit what it can access, and limit what it can finish without a human.

Diagram showing an AI agent action request passing through allowed-action list, scoped permissions, reversibility check, human approval, execution logging, and blast-radius limits.
Give the agent the smallest surface that works
Grant the agent the minimum set of tools and the minimum access its job needs, and nothing beyond that. This is the principle of least privilege, defined by NIST as restricting access "to the minimum necessary to accomplish assigned tasks," applied to a non-human actor. An agent wired to a human administrator's credentials inherits every door that human can open; an agent given a narrow, purpose-scoped credential can only reach what its task requires.
Restrict the agent to an allowed-action list
Give the agent an explicit list of tools it may call, and refuse anything outside it by default. OWASP's guidance on excessive agency is to "limit the extensions that LLM agents are allowed to call to only the minimum necessary." An allowlist turns tool access into a deliberate decision rather than a side effect of what the model can reach, and it means a prompt that talks the agent into trying something new fails closed instead of open.
Put the real limits in the credentials
Telling the agent in its prompt to be careful is not a control; the permissions attached to its credentials are. Scope them the way you would for a narrow service account:
- Read-only access wherever the agent only needs to read.
- Write access limited to the specific records, tables, or endpoints in its remit.
- A separate identity per agent, so its actions are attributable and its access can be revoked on its own.
Scoping credentials is ordinary access-control work, and it is where identity and access controls from a security-minded team pay off. The model may decide what to attempt; the permissions decide what actually goes through.
Putting an acting agent near production data? Map its failure modes with our engineers before it gets write access. A short review of the tools, permissions, and approval points usually surfaces the one-way doors that need a gate, while the agent is still cheap to change
Put a human on the one-way doors
Require human approval on the actions that cannot be undone, and let everything reversible run unattended. This is the human-in-the-loop control OWASP names directly — "require a human to approve high-impact actions" — and the reason to scope it tightly is that approval fatigue is its own failure. If every step needs a click, reviewers start rubber-stamping, and the gate becomes theatre. Reserve it for the one-way doors, and it stays meaningful.
Design the approval step so a human can actually judge it: show what the agent intends to do, the inputs it used, and the reason it gave, in one place, with a single approve or reject. Anthropic's write-up on building effective agents frames the same idea as pausing "for human feedback at checkpoints," alongside stopping conditions like a maximum number of iterations so a run cannot escalate on its own. The aim is a fast, well-framed decision rather than a queue people learn to clear without reading.
Bound the blast radius
Assume a wrong action will get through eventually, and cap what it can do when it does. Blast radius is the amount of damage a single mistake can cause before something stops it, and shrinking it is what turns an incident into a footnote. Most of the levers are unglamorous and effective:
- Caps and rate limits — a refund ceiling, a maximum number of actions per run and per day, so a loop cannot drain an account.
- Sandbox and dry-run — destructive actions run against a test target or return a preview first, so the real system is touched only after a check.
- Reversibility by default — soft-delete instead of hard-delete, and idempotent actions so a retry does not fire the same payment twice.
These limits also contain the failure that agents are prone to when real input gets messy: looping and repeating an action. Caps and idempotency keep that contained, and a companion article in this series covers the timeouts, retries, and cost caps that stop runaway runs in more depth (PENDING — internal link to follow). The pattern throughout is the same as the rest of AI agent development done well: expect the wrong action, and make it cheap.
Make every action traceable
Log every action the agent takes with its trigger, its inputs, and the reason the agent gave for it. Without that record, a wrong action is both unexplainable and hard to undo, because no one can see what the agent thought it was doing or what it touched. With it, an incident becomes a specific, reversible event: you can find the action, understand the decision, and roll it back.
Traceability is also what rebuilds trust after the first mistake. A team that can point to exactly what happened, why, and how it was contained will keep an agent running; a team left guessing will switch it off, which is how most agent projects quietly end. The log is not paperwork — it is the thing that lets the agent stay in production after it errs.
Key takeaways
- An acting agent's risk is set by the worst single action it can take unsupervised, so average accuracy is the wrong yardstick.
- Decide autonomy and authority separately: an agent can run many steps on its own and still need a human before any irreversible one.
- Apply least privilege — an allowed-action list and scoped credentials — so the agent can only reach what its task requires.
- Reserve human approval for the one-way doors, or approval fatigue turns the gate into rubber-stamping.
- Bound the blast radius with caps, sandboxes, and reversible actions, and log everything so a mistake is visible and undoable.
Why the guardrails decide whether the agent survives its first mistake
An acting agent will get something wrong eventually, and whether that ends the project depends entirely on what was built around it. With the worst action contained, a mistake is a logged, reversible event that the team fixes and moves past. Without those controls, the same mistake is an incident with a customer on the other end, and the agent goes back in the box regardless of how useful it was the rest of the time.
The work is deciding these limits on purpose — what the agent may do, what it may reach, what needs a human, and how far any one action can go — before it touches production rather than after an incident.
If you are putting an agent near real systems, a dedicated engineering team can design the permission model, approval gates, and blast-radius limits alongside your own engineers, so the agent ships with its guardrails rather than acquiring them the hard way.
Summarize with AI
FAQ
Interesting For You

AI testing in EdTech: how to catch regressions before launch
The prompt change fixed the support assistant’s bad answer on Tuesday; by Friday, the same release had broken citation accuracy in cases nobody retested. AI testing in EdTech has to catch that kind of regression before users discover it. A conventional test suite can confirm that APIs return 200 responses and schemas remain valid. The generated answer can still become less accurate, less complete, or less grounded after a model, prompt, retrieval, or data-source change.
Read article

Confident, wrong answers: fixing retrieval before you blame the model
A wrong answer from an AI assistant would be easy to catch if it looked unsure, and it never does. Three colleagues forward the same screenshot: the internal assistant answered a policy question with complete confidence, and it was wrong. After a few of those, people stop asking it and go back to messaging each other.
Read article

AI document ingestion in EdTech: what breaks first
Education software receives institutional policies, faculty handbooks, admissions records, support knowledge, assessment material, administrative forms, and user-uploaded files. The key engineering questions are where structure can be lost, which failures should stop processing, and which checks belong in deterministic code before an LLM is called. Those decisions determine whether the feature remains debuggable when clean demo files give way to real inputs.
Read article


