
Explainable AI Made Judgment Worse
A field experiment found AI explanations made evaluators worse at disagreeing with the machine, which breaks the core assumption behind human oversight.
Every AI governance framework I have read makes the same bet: if the machine explains itself, the human can check it. The explanation is the bridge between the model's output and the reviewer's judgment. A field experiment published in Harvard Business Review last week tested that bet directly. The explanation was not the bridge. It was the anaesthetic.
What the experiment found
Leonid Sudakov and Nathan Furr had 228 experienced evaluators assess 48 submissions to an MIT global social impact innovation challenge, under three conditions: human-only evaluation, an LLM's evaluation accompanied by a narrative explanation of its reasoning, and an LLM's evaluation with no explanation at all. Everything was scored against the assessments of four expert judges.
Evaluators accepted the LLM's recommendation 67% of the time. They agreed with the LLM's decisions roughly 75% of the time whether or not it explained itself, against 54% agreement with decisions made by other humans. And the narrative explanations made one thing substantially worse: false negatives. Evaluators who read the model's fluent reasoning rejected viable ideas the experts would have kept alive. People did better when they were given no reason at all for the AI's decision.
The researchers put it plainly: "Our findings reveal that LLM explanations do not necessarily improve decision-making. Effective human-AI collaboration requires designs that preserve rather than supplant independent human judgment."
There is a reasonable objection here, and I want to put it properly: maybe deferring to a strong model is simply rational. If the machine is usually right, agreeing with it is not a failure of judgment, it is good judgment. But that reading does not survive the numbers. If the explanations were doing any checking work, agreement should differ between the condition with reasoning and the condition without. It did not. Agreement sat at 75% either way, which means nobody was evaluating the argument. And where the explanation did move behaviour, it moved it the wrong way: more good ideas killed. A confident, reasoned case is harder to argue with than a bare verdict. That is true of people, and it turns out to be true of machines. The explanation does not open the output for inspection. It closes the case.
Oversight is a capability, not a control
Article 14 of the EU AI Act asks that the people overseeing a high-risk system "remain aware of the possible tendency of automatically relying or over-relying on the output" and that they be able to "disregard, override or reverse" it. The enforcement date for stand-alone high-risk systems moved to December 2027; I covered the timing in Monday's special edition. The obligation itself has not moved. Read those two phrases again. Neither describes a button. Both describe something a person has to be able to do: notice their own automation bias while it is happening, and hold their own read of a case long enough to disagree with a fluent machine.
That is a capability. And the review step most organisations have built trains the opposite of it. The human sees the model's answer first, then the model's reasoning, then a field asking whether they agree. Every element of that design makes independent judgment harder to form, because by the time the reviewer starts thinking, the anchor is already set and the argument for it already made. What worries me is not that people defer to machines. It is that we are designing the deference in deliberately, labelling it transparency, and reporting it upward as human oversight.
Change the order, not the tooling
There is already a design answer, and it is cheap. Zana Buçinca, Krzysztof Gajos and Maja Malaya built a decision-support system that uses reinforcement learning to decide, case by case, how much help to give: a full recommendation, a partial explanation, or nothing at all. In two experiments with 316 and 964 participants, the adaptive system beat standard uniform AI support on decision accuracy, and frequently the human-AI pair outperformed either the human alone or the AI alone. The blunt version: sometimes the best thing the system can do is withhold its answer.
You do not need reinforcement learning to use that finding. You need to change a sequence. Pick one recurring decision in your organisation where a model already produces a recommendation and a human already reviews it. Grant applications, intake triage, candidate screening, whatever runs weekly. From this Thursday, the reviewer writes their own one-line call first, before opening the AI's output. Then they open it, compare, and log every case where the two differed. No new tool, no budget line, no committee. By Friday you have something almost nobody has: a record of where human judgment and the model actually diverge, which is exactly the evidence Article 14 assumes you can produce.
If your reviewers never disagree with the model, you do not have human oversight. You have a very expensive agree button.
Sources
Your move
See where your organisation actually stands.
The free VERIFY capability scan scores you across the six moves in ten minutes.
Get the playbook