The Case for Risk-Based Agent Permissions

post-thumb

An agent that drafts a refund email is doing something very different from an agent that sends one. The first puts words on a screen, and a person reads them before anything happens. The second moves money and lands in a customer’s inbox, where you can’t pull it back. Same task, same text, very different stakes.

So the question “what should an agent be allowed to do?” usually gets a worse answer than it deserves. The common answer sorts permission by stage. Reading is safe, drafting is safe, acting is risky. That is close enough to be wrong. Risk should decide an agent’s permissions. How much damage can the action cause? How easily can you undo it? How sensitive is the data it touches? And how well do you understand the risk to begin with? Stage only hints at those answers.

Two camps make me impatient. One wants full autonomy now and treats guardrails as friction. The other approves every click and calls it safety. Both are lazy, just in opposite directions.

Stage is a weaker signal than it looks

Agents can do a wide range of things. Think of it as a ladder:

  • Read-only research. The agent gathers and reads. It pulls data, searches sources, and summarizes. It changes nothing.
  • Drafting and staging. The agent produces something, like an email, a report, or a proposed change. The output sits and waits for a person.
  • Recommendations. The agent tells you what it would do next and offers it as a choice you click to confirm.
  • Approved execution. The agent takes a real action in a connected system. It sends the email, updates the record, or issues the refund.
  • Bounded standing authority. The agent acts on its own inside fixed limits, and anything past those limits stops for a human.

The Cloud Security Alliance maps a similar climb. It runs from no autonomy and assisted action through supervised plans and bounded conditional action, up to full autonomy. It is blunt that full enterprise autonomy lacks adequate controls today.

Stage alone misleads you for two reasons. Reading isn’t automatically safe. An agent with read-only access to a payroll system can still leak data if it can send that data somewhere. And acting isn’t automatically dangerous. An agent that re-tags support tickets takes an action, but a wrong tag takes seconds to fix. The stage of a task doesn’t tell you the size of the mistake.

Hold two examples in mind. The low-stakes one is an agent that compiles a weekly competitor summary from public web pages. It reads, it drafts, and a person skims the result. If it gets something wrong, you notice and move on. The high-stakes one is an agent that issues customer refunds and emails the customer. Wrong amount, wrong account, or wrong tone, and you’re clawing back money and apologizing.

Why “approve every action” quietly fails

Putting a human in the loop for every step is a sound instinct. In practice, it falls apart fast.

Anthropic looked at how people actually use agents. It found that most tool calls still run with safeguards and a human in the loop, but that oversight drops as people get more comfortable. That matches what security researchers keep seeing. When agents ask for approval constantly, people approve at very high rates without reading closely. The name for this is approval fatigue, and it is the main way per-action approval breaks down.

Think about what you’re really asking a person to do. Click approve forty times a day on actions that are almost always fine, then catch the one that isn’t. Nobody keeps that up. You’ve built a process that looks careful and behaves carelessly.

So per-action approval is the wrong default for anything routine. It trains people to rubber-stamp, and a rubber stamp is worse than no check at all. It leaves a paper trail saying someone checked when nobody did. We have covered the same failure in The Approval Bottleneck That’s Quietly Killing Your AI Agent ROI.

The fix is to move the human decision up a level. Approve the plan, not each keystroke. A person confirms “yes, refund these twelve customers for this reason, up to this amount each,” and the agent carries out that plan. Save the interruptions for anything outside the plan or over a risk line. The Frontier Model Forum argues for this same shape. Govern at the level of the workflow, and escalate by risk tier instead of treating every action the same.

So how do you decide what to allow?

Score the action, not the stage. Three questions do most of the work.

How bad is the worst outcome? A wrong competitor summary wastes ten minutes. A wrong refund costs money and trust. Those belong in different buckets.

How hard is it to undo? This is reversibility, and it matters more than raw size. Singapore’s IMDA guidance recommends sorting agent tasks into tiers by severity and reversibility. A draft is fully reversible. A tag nearly is. A sent email and a processed payment aren’t.

How well do you understand the risk? People skip this one. Writing in Harvard Business Review, Mike Walsh argues that the right level of autonomy depends on how well you understand the risk, and not only on how big it is. A large risk you understand well can be bounded. A small risk you don’t understand yet should wait, because the failure that hurts you is usually the one you didn’t model. MIT Sloan makes the same point: assess risk and business value deliberately before you hand over control.

Put those together and the competitor summary agent can run on a long leash with a quick glance. The refund agent can’t, however reliable it looks on a good day.

Give an agent the least power that does the job

Once you know the risk, shape the permissions to match it. The principle is least privilege. Grant the smallest set of capabilities that lets the agent finish the work, and nothing extra.

On Autohive, you choose which capabilities an agent can use. Each capability can be switched on or off with a click, so the agent only gets access to the tools needed for its role. You can give a research agent read-only capabilities while keeping write actions switched off. It can search the web and draft a competitor summary, but it cannot send anything outbound unless you choose to give it that capability.

Two more ideas make least privilege real. The first is delegated identity. The agent acts on behalf of a specific person, using that person’s authorized access instead of an all-powerful service account. The second is scope attenuation, which means an agent can narrow its own authority but never widen it. The OpenID Foundation’s work on agent identity lays out the machinery, from OAuth and scoped permissions to keeping the system that decides policy separate from the system that enforces it. IMDA states the human principle plainly: an agent should never inherit more authority than the person who granted it.

One combination deserves a hard line. The Frontier Model Forum warns against building a single high-risk agent that can touch private data, read untrusted outside content, and send messages to the world all at once. That mix is how a poisoned web page turns into leaked customer data. If you need all three functions, split them across separate agents with separate rights, and pass work between them. A research agent with no send rights can hand findings to a publishing agent that never sees the raw private data.

Controls that actually hold

Permissions decide what an agent can do. Controls decide whether you can trust it, check it, and stop it. The useful ones aren’t exotic:

  • Policy rules. Encode limits as rules the system enforces, not guidelines you hope people remember. “Refunds over $500 always escalate” should be code, not culture.
  • Pre-execution gates. For consequential actions, record what the agent is about to do before it does it, in a log it can’t quietly edit. Emerging audit guidance pushes for tamper-evident records written before execution and kept independently of the agent.
  • Audit logs. Keep a clear history of who did what, agent or human, so you can reconstruct any decision after the fact. Our AI Agent Audit Trail guide explains what that record needs to contain.
  • Revocation. You need a fast off switch. Cut a session, pull an integration, or stop an agent mid-run without redeploying anything.
  • Named ownership. Every agent with real authority has a specific human accountable for it. “The system approved it” is not an answer anyone should accept.

Autohive supports a lot of this scaffolding: role-based access, audit logs, session revocation, encryption, and workspace controls that keep raw source data in read-only directories so an agent can’t alter its own inputs. The security and compliance docs cover the details. When an agent lacks authorization for something, it doesn’t fail silently or route around the limit. It raises a connection prompt in the conversation and waits for a person. That is the behavior you want: blocked by default, unblocked on purpose. The same logic covers scheduled recurring jobs. They’re fine when the task is well understood and reversible, and they write to the same logs as everything else.

A rule you can actually use

Here is the rule I’d hand a team today.

Before you let an agent act on its own, score the action on three things: how bad the worst outcome is, how hard it is to reverse, and how well you understand what could go wrong.

If the worst case is small and you can undo it, let the agent act and log it. The competitor summary lives here.

If you can undo it but the blast radius is wide, let the agent act within fixed limits and escalate anything past them. Small refunds to verified customers, up to a capped amount, with a plan a person approved once, live here.

If you can’t undo it, or you don’t yet understand the risk, it stops for a named human who approves the specific plan, not each click. Large refunds, new refund types, and anything touching an account you haven’t seen before live here.

Re-score as you learn. An action that starts in the top tier can move down once you understand its failure modes and have the logs to prove it behaves. That movement, earned instead of assumed, is what responsible autonomy looks like.

Back to the refund email. Drafting it was never the risky part. Treating drafting and sending as two points on a safety line misses the real question. That question is what happens when the agent is wrong, how fast you find out, and who answers for it. Decide that first and the permissions sort themselves out.

If you want a practical next step, our guide on how to build a custom agent walks through assigning capabilities one at a time, which is where every one of these decisions actually gets made.

You may also like