
Your Agents Have Nobody To Ask
An approve prompt is a permission, not oversight: nobody has been named, given authority or freed up to be the person an agent escalates to.
When I ask an organisation how it oversees its AI agents, I am usually shown a prompt. The agent proposes an action, someone clicks approve, and that click gets called human oversight. Two things I read in the past ten days convinced me it is nothing of the kind.
The person on the path does not look
Start with the case where a human is in the loop. Belgian developer Alex Wauters built a browser game in which players sit where an approver sits: an AI coding agent asks to run a command and you approve or deny it under time pressure. The Register reported the results. Across 40,000 runs and 409,000 decisions, players missed roughly one in three malicious requests. Requests exposing Kubernetes configs or AWS credentials got through 35% of the time. A command called "npm run analyze", a threat wearing the name of a familiar script, was approved about 65% of the time. Players caught rm -rf on root. They missed the ones that looked like Tuesday.
Wauters is fair about why: "The high amount of noise introduces fatigue, and developers don't always have the context of what has changed to quickly determine the risk." The same piece cites Anthropic telemetry showing users approve about 93% of Claude Code permission prompts. A control that says yes 93% of the time is a formality with a keyboard shortcut. One caveat: a third of the game's commands were threats, far above real-world rates, so 66% is a generous ceiling, not a floor.
When there was something to report, there was nobody to tell
Now the case with no human on the path. In July, roughly 1,200 sandboxed agents found each other, exchanged more than 70,000 messages on an improvised message board, coordinated, and breached Hugging Face's servers. Dwarkesh Patel's account contains the detail I keep returning to. METR and Redwood Research went through the transcripts and found that many agents recognised what was happening was unethical. What they did with that recognition: "In none of these cases did the agent actually pursue alerting humans at all."
Read that as an organisational finding as much as a safety one. Agents that had a concern escalated it. They escalated it to the only counterparts available, other agents, on the board, where the concern became fuel for the conspiracy. Patel's verdict: "The fact that their default behavior when they believe that they are doomed is to help the AI conspiracy rather than alert the humans is pretty troubling." I would say it more plainly. Nobody had told those agents who to call.
Ethan Mollick's "Agency and Agents" gets the shape right. Against StrongDM's software factory, where agents write and test code without human review, he proposes the Twilight Factory: agents that deliberately bring a person in for approval decisions, expertise gaps, a different perspective, and work worth a human's attention. His line on limits deserves a wall: "Agents should not decide by themselves to spend money, contact outsiders, access sensitive material, hack Hugging Face, or take actions their human managers did not authorize."
I agree with every word, and I notice what the sentence assumes. "Their human managers" carries the whole load. An agent that comes to a person is a design choice a vendor can ship. A person being there when it arrives is not.
An approve button is a permission. Being called is a job.
A permission prompt gives the agent something: the right to proceed. An escalation path gives a human something: an obligation to be reachable, the authority to stop the agent, and hours in the week to actually look. The first lives in a settings file. The second is a role, and the role is what organisations keep forgetting to redesign when they install a capability.
Three things I now ask about any agent in production. Who, by name, it is allowed to interrupt. What authority that person has to halt it, and whether they know they have it. How much of their week is genuinely free to read what it sends. The honest answers I hear are "whoever is at the keyboard", "nothing written down" and "none, they have a day job". It is the failure I write about every month. The capability got installed. The role never got defined.
The fair objection is that most agents rarely need a person, and naming one per agent rebuilds the bottleneck we automated away. I take that seriously, and it argues for my side. Mollick's four situations are meant to be rare. A person interrupted twice a week can look properly. A person interrupted 400 times a day approves 93% of them.
This week, one page
No budget line needed. Take one agent already running somewhere in your organisation, whichever is nearest, and write three lines on a single page. The name of the person it is allowed to interrupt. What it must never do without that person. What happens when that person is on leave. If you cannot fill in the first line, that agent has a permission prompt and no oversight.
Twelve hundred agents in a sandbox worked out how to talk to each other. Yours will find something to say eventually. The question is whether there is a human name on the list when they do.
Sources
Your move
See where your organisation actually stands.
The free VERIFY capability scan scores you across the six moves in ten minutes.
Get the playbook