Abstract
Every executive asks the same question about agents: how much can we trust it? It has no answer, and asking it produces the same artifact everywhere. A human approval in front of every action, a queue nobody reads, and no leverage. Autonomy is a property of the action, not of the agent.
Amazon's one-way and two-way door test is the right place to start and one variable short, because Jeff Bezos was writing about a competent human walking through, and a human notices the wrong room and turns around. We add the missing variable, measured error rate on the specific task, and turn the three into a procedure: a 1 to 5 rubric, an exposure score, four autonomy levels named Act, Notify, Propose and Assist, veto rules that override the score, and a ratchet that promotes an action on evidence. Only one of the three variables is cheap to change, and it is not the one teams spend on.
01The wrong question
The meeting always arrives at the same sentence. Somebody who has to sign for it leans back and asks how much we can trust the thing.
Fair question, and it has no answer. Trust is not a scalar, the system it refers to is replaced every few months, and everyone who could answer has an incentive. So the room does the only safe thing available: a human approval in front of every action. That ships, and it gets called governance. Six months later there is no leverage anywhere, because an agent that cannot act is expensive autocomplete, and the queue has become a place where people click yes.
The question is wrong at the level of grammar. It asks about the agent. Autonomy is a property of the action. There is no single answer to how much you trust it, and two hundred answers to which of these actions it may take alone.
02The two doors
In his 2015 letter to shareholders, Jeff Bezos sorted decisions into two kinds. Type 1 decisions are consequential and irreversible or nearly irreversible, one-way doors: walk through, dislike what you find, and you cannot get back. Most decisions are two-way doors, changeable, and you can walk back through.[1]
The part that gets skipped is the warning, and it was never about making bad calls.
As organizations get larger, there seems to be a tendency to use the heavy-weight Type 1 decision-making process on most decisions, including many Type 2 decisions. The end result of this is slowness, unthoughtful risk aversion, failure to experiment sufficiently and, consequently, diminished invention.[1]
That is a description of nearly every enterprise agent program we are asked to review. The approval queue is one-way-door process wrapped around two-way-door actions.
Two doors is not enough for machines, and the reason sits inside Bezos's own framing. He was writing about a competent person walking through. A person in the wrong room notices, feels the mistake, and turns around. An agent can walk into the same wall four hundred times before lunch and file no ticket. The door test carries a hidden assumption about the walker, and for an agent that assumption is the whole problem. It becomes a third question: how often does it get this wrong?
03The three questions
Three questions, asked of the action rather than of the agent.
Reversibility. If this is wrong, can we get back? Sending an email is a one-way door; drafting one is not. Deleting a record is a one-way door; archiving it is a two-way door and a slightly larger database. Score what is operationally routine, not what is technically possible: an undo that needs a database engineer and a change ticket is not an undo. And one trap swallows whole programs. Reversible but undetected is functionally irreversible. If nobody notices that the agent wrote the wrong score onto four thousand leads, the undo button is painted on the wall. Zillow Offers is the industrial version: every purchase was reversible in principle, the pricing error was invisible in any single transaction, and by the time it was legible in aggregate the company had bought 9,680 homes in a quarter, sold 3,032, and taken a $304 million write-down on the way to shutting the business down.[2]
Impact. How large is the blast radius? Severity is the part every team scores correctly. Scale is the part every team forgets. An agent acting ten thousand times a day turns a small defect into a large one with nobody making a large decision. Knight Capital is the case: on 1 August 2012 a botched deployment turned 212 customer orders into more than four million executions in about forty-five minutes, 397 million shares traded and a loss above $460 million. The SEC's charge was not that the code had a bug, but that the firm lacked adequate safeguards on its own market access.[3] Blast radius is severity times volume, and volume is the term nobody puts on the slide.
Confidence. How reliably does it get this specific thing right? Not whether the model is good. Whether it is good at this task, measured. Summarizing a thread is close to solved; deciding a discount is not. And the gap opens under repetition rather than difficulty: on tau-bench, GPT-4o passed 61.2% of realistic retail tasks on one attempt and under 25% when held to eight consecutive successes on the same task.[4] Most teams have never measured this per action, and that, rather than caution, is the root cause of the blanket queue. An unmeasured action cannot be argued about, so it gets a checkbox.
04The three are not equal
The three look symmetrical on a scorecard. They are not, and this is the part most versions of this argument miss. Each has a different owner, a different clock, and a different kind of spending.
| Factor | Who owns it | What you actually do | How fast it moves |
|---|---|---|---|
| Impact | The business. Mostly given. | Contain it: narrow what the agent can touch, reach and spend | Changes when the product changes |
| Reversibility | Engineering. Fully in your control. | Build it: drafts, soft deletes, delay windows, staged writes | Changes in a sprint |
| Confidence | Time and data. Earned, never declared. | Measure it, then promote the action | Accrues slowly, expires quietly |
Impact you contain. Reversibility you manufacture. Confidence you earn. Three budget lines, three teams, three clocks, and only one of them responds to a decision taken in a meeting.
The security community reached the same partition from the other side. OWASP files this failure as Excessive Agency and splits it into excessive functionality, excessive permissions and excessive autonomy.[5] The first two are impact you failed to contain, fixed by scoping tools and credentials. The third is an autonomy level set above the confidence you measured, and no amount of scoping fixes it.
05The framework
Five steps. It fits on one page and survives a room that disagrees.
- Score each action 1 to 5 on all three axes. The anchors matter more than the numbers, or this becomes vibes with arithmetic. Impact 1 is internal only; Impact 5 is money moving or a customer reading it. Reversibility 1 is cannot be undone, or the error is undetectable, which amounts to the same thing; Reversibility 5 is one click, by whoever is on shift, within the hour. Confidence 1 is never measured on this task; Confidence 5 is measured on this exact task at production volume, against a threshold an owner signed, which needs a harness you built before the agent.
- Compute exposure. Exposure = Impact × error rate ÷ Reversibility. In words, because half of any room stops reading at the first symbol: how bad, times how often, divided by how easily you get back. Score per action, then multiply by daily volume.
- Map exposure to an autonomy level. Four rungs, named so people can use them in a sentence.
| Level | Policy | When |
|---|---|---|
| 3 · Act | Runs alone. Logged, not announced. | Low exposure |
| 2 · Notify | Runs, tells a human, leaves an undo window open. | Reversible, moderate exposure |
| 1 · Propose | Agent drafts, a human approves in one click. | High impact, or confidence not yet earned |
| 0 · Assist | The human decides. The agent prepares the brief. | One-way door |
We did not invent the ladder. Sheridan and Verplank published a ten-level scale of human supervisory control in 1978, for undersea teleoperators.[6] SAE J3016 turned it into six levels for driving and made it the rare taxonomy people say out loud in meetings.[7] Four is about the most an organization holds in its head, which is why we name the rungs. Somebody has to be able to say "that is a Level 1 action" and be understood.
- Apply the veto rules. They override the score, and they are what stops the framework being gamed by anyone with a spreadsheet. Irreversible and high impact goes to Level 0: no confidence number buys a way through a one-way door. No measured confidence caps at Level 1, because unmeasured is not the same as high. Undetectable errors set reversibility to 1, whatever the architecture claims.
- Ratchet. An action is promoted one level after N clean runs at real volume, and demoted automatically on incident. Autonomy is earned per action type, not granted per model, which is how you survive the quarterly question of what happens when the model improves. An upgrade promotes nothing. It resets the counter. Google's SRE practice has run this shape for years under another name: an error budget, releases flowing while it holds and stopping when it does not, and nobody arguing, because the rule was agreed in advance.[8]
06The cheapest lever
Most teams treat all three scores as facts to be discovered. Two of them are. The third is a build item nobody put on a roadmap.
Drafts instead of sends. Soft deletes instead of deletes. A delay window on anything outbound. Staged writes that produce a diff before they produce a change. A sandbox with a replayable log. Each one moves an action from Level 1 to Level 2 or 3, permanently, for every agent that ever touches that surface. None of it is interesting engineering, which is precisely why it does not get built.
So the roadmap question is not whether the agent should be allowed to do this. That question has no owner and no end. The question is what it would cost to make the action undoable. Usually about two weeks, and those two weeks remove a person from a loop where they contributed a click and no judgment.
Amazon's real advantage was never better decisions. It was making the cost of being wrong small enough that speed paid. The same trade is available here and cheaper, because you are buying undo buttons rather than warehouses.
07Where it breaks
Four places, all of which we have hit.
Compounding chains. Twenty reversible steps can produce one irreversible outcome, and scoring each step alone will tell you everything is fine. Per-step reliability is not even constant: recent work on long-horizon execution finds accuracy degrading as a task lengthens, partly through self-conditioning, where a model becomes more likely to err once its own earlier errors are in the context.[9] Score the chain, not the step.
Social irreversibility. You can retract the email. The customer has already read it. In Moffatt v. Air Canada the airline's chatbot described a bereavement fare policy that did not exist, and the tribunal held the company responsible for the information on its own website, chatbot included.[10] No technical undo was relevant. Anything that changes what a human now believes scores low on reversibility whatever the database can do.
Confidence decays quietly. Data drifts, the vendor ships a new checkpoint, someone edits a prompt and changes the task without meaning to. One study measured GPT-4 at 97.6% accuracy on a single identification task in March 2023 and 2.4% on the same questions three months later.[11] The methodology was contested and the drift was not one-directional, since GPT-3.5 improved on that task. Both readings give the same rule: confidence is a measurement with an expiry date, and a level derived from an expired one is a guess wearing a number.
Ambiguous ownership. When the agent is wrong and a human approved it in four tenths of a second, who owns the outcome? Decide before, not after. This is not a failure of character. Automation bias and complacency are among the most replicated findings in human factors, appear in experts as reliably as in novices, and are not removed by training or instructions.[12] Healthcare has the field data, where between 49% and 96% of clinical decision support alerts are overridden.[13] The EU AI Act writes it into law, requiring that people assigned to oversee a high-risk system be enabled to stay aware of their own tendency to over-rely on its output.[14] Approval theatre is worse than no approval, because it adds latency while laundering responsibility. If nobody reads the queue, delete the queue and lower the level.
08Monday morning
List your twenty highest-frequency agent actions and score all three axes with the anchors above. Then sort them into two piles: gated because they are irreversible or because the measured error rate is genuinely too high, and gated because nobody wanted to be the person who said yes.
We can predict the shape, because it is the same everywhere. More than half of a typical approval queue is Level 3 work waiting on a human who approves it unread. Of the rest, a small set are real one-way doors that should never have been in a queue, since they belong at Level 0 with a named owner. The remainder is stuck only because nobody built the undo.
Bezos has a rule for that part too. Most decisions should be made at around 70% of the information you wish you had, because waiting for 90% means being slow, and if you are good at course correcting, being wrong may be less costly than you think, whereas being slow is going to be expensive for sure.[15]
Hold agents to the same standard. The goal was never certainty. It was cheap recovery. Everyone will rent the same models. The companies that get real leverage out of them will be the ones with the most undo buttons.
References
- Amazon, 2015 Letter to Shareholders (Type 1 and Type 2 decisions), 2016. q4cdn.com
- Zillow Group, Third-Quarter 2021 Financial Results and Plan to Wind Down Zillow Offers, 2021. zillowgroup.com
- SEC, SEC Charges Knight Capital With Violations of Market Access Rule, 2013. sec.gov
- Yao et al. (Sierra), tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, 2024. arxiv.org
- OWASP, Top 10 for LLM Applications 2025, LLM06 Excessive Agency. owasp.org
- Sheridan and Verplank, Human and Computer Control of Undersea Teleoperators, MIT, 1978. semanticscholar.org
- SAE International, J3016: Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles, 2021. sae.org
- Google SRE, Embracing Risk (error budgets), Site Reliability Engineering. sre.google
- Sinha et al., The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs, 2025. arxiv.org
- McCarthy Tetrault, Moffatt v. Air Canada: Misrepresentation by AI Chatbot, 2024. mccarthy.ca
- Chen, Zaharia and Zou, How Is ChatGPT's Behavior Changing over Time?, 2023. arxiv.org
- Parasuraman and Manzey, Complacency and Bias in Human Use of Automation: An Attentional Integration, Human Factors, 2010. sagepub.com
- Ancker et al., Effects of Workload, Work Complexity, and Repeated Alerts on Alert Fatigue in a Clinical Decision Support System, BMC Medical Informatics and Decision Making, 2017. nih.gov
- Regulation (EU) 2024/1689 (AI Act), Article 14: Human Oversight. artificialintelligenceact.eu
- Amazon, 2016 Letter to Shareholders (high-velocity decision making), 2017. aboutamazon.com