Image created using AI Tools
Hook
I attended an AI expo last week.
I spoke with a lot of founders and executives, and almost every one of them told me some version of the same thing. Their system learns from its mistakes. It coaches the humans who use it. It watches its own performance and adjusts. A human described retraining his model and told me that afterward it automatically goes back and rescores everything.
I nodded and moved to the next booth, and the sentence stayed with me for the rest of the day.
In ten years of building and reviewing these systems, I have almost never seen an AI correct itself. I have seen plenty of systems change. A model gets retrained, a prompt gets edited, a threshold gets moved, a reference dataset gets refreshed. Something is different afterward. But changing is not correcting. Correcting requires knowing you were wrong, and knowing you were wrong requires a reference outside yourself.
When a model rescores its old decisions using its own new weights, it is not checking its work. It is restating its opinion with more confidence. If it was wrong before, the rescore agrees with the error. The only thing that has changed is that the disagreement has disappeared, and disagreement was the only signal anyone had.
I am not particularly troubled by the claim itself. I am troubled by what follows from it if it is wrong.
Stakes
Here is why this matters more than it looks.
Almost every safeguard anyone has built around AI rests on one control, and that control is a person. The EU AI Act requires human oversight for high risk systems. The NIST AI Risk Management Framework returns to human review at multiple points. Medicare Advantage rules require a physician to review before an adverse coverage determination issues. Every vendor deck I have read in the last two years has a slide with a small human icon on it and the words human in the loop.
The phrase does an enormous amount of work. It is what allows a hospital to deploy a diagnostic aid, a payer to run automated review, a bank to score applications, and a company to put an agent in front of customers. In each case the argument is the same. The machine proposes and a qualified person disposes, so the accountability that has always sat with the organization still sits there, undisturbed.
That argument is only as good as the person. And the person is only a control if they do something. A reviewer who reads every recommendation and approves every recommendation is not a safeguard. They are a rubber stamp with a license number attached.
We have a rough sense of how often this happens, and the numbers are not reassuring. Across eleven studies covering 570,776 prescriptions, clinicians overrode drug interaction alerts about ninety percent of the time. That statistic is usually told as a story about alert fatigue, which it partly is. But read it the other way and it says something harder. In a workflow explicitly designed so a human checks the machine, the human and the machine disagree almost always, and nobody in that literature can tell you how often the human was right to disagree.
If the loop is hollow, then a great deal of what has been built on top of it is hollow too. Not fraudulently. Just hollow.
Reversal
The obvious response is to say that the humans need to try harder. Train the reviewers. Reduce the alerts so the ones that remain matter. Give people time. Make the interface better. Put the right person in the seat.
I am not saying that training and interface design do not matter, because they clearly do. I am saying that the evidence does not support the assumption underneath the obvious response, which is that adding a human to an automated decision makes the decision better.
In 2024 Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone published a preregistered systematic review and meta-analysis in Nature Human Behavior covering 106 experimental studies and 370 effect sizes. On average, human and AI combinations performed significantly worse than the better of the two alone. Not worse than the human. Not worse than the machine. Worse than whichever of them was better at the task.
The detail underneath is still sharper. They found performance gains when the combination involved content creation, and performance losses when it involved making decisions. And they found that when the human was better than the AI, combining them helped, while when the AI was better than the human, combining them hurt.
Read that last sentence slowly, because it inverts the standard justification. We add human oversight precisely in the settings where we have deployed a machine that we believe outperforms unaided human judgment. That is the condition under which the meta-analysis says the combination loses.
So the problem is not that organizations put weak humans in the loop. The problem is that putting a human in the loop is not, on the evidence, a reliable way to make a decision better, and we have been treating it as though it were self-evidently one.
Autopsy
There are at least four separate failures hiding inside the single phrase human in the loop, and they deserve different answers.
The first failure is that people defer.
Forty years ago, Lisanne Bainbridge described the ironies of automation: the more reliable you make an automated system, the less practice the human operator gets at the task, and the worse equipped they are for the moment the automation fails. Automation does not remove the human contribution to a system. It moves it, usually to the hardest part, and then takes away the conditions that would have kept the human sharp.
Kate Goddard, Abdul Roudsari and Jeremy Wyatt reviewed automation bias in clinical decision support in 2012 and found it was common rather than exceptional. David Lyell and Enrico Coiera followed with a review connecting automation bias to verification complexity: the harder it is for a person to independently check the machine’s recommendation, the more likely they are to accept it. Raja Parasuraman and Dietrich Manzey described the same thing from the attention side, as complacency.
There is a particular study I keep reading over and again. In 2004 Eugenio Alberdi and colleagues looked at what happened when computer aided detection gave incorrect output to radiologists reading mammograms. The wrong prompt did not simply fail to help. It changed what the reader saw. The safeguard became a source of error.
The second failure is that people defer selectively.
This one is worse than uniform deference, and it is the finding that should trouble anyone who has used human review as a fairness argument.
Saar Alon-Barkat and Madalina Busuioc ran three experiments in the Netherlands and published the results in the Journal of Public Administration Research and Theory in 2023. They were looking for automation bias and they found something else. Participants did not follow the algorithm uniformly. They followed it when its advice matched what they already believed. When the algorithm gave a low score to a teacher with a Moroccan sounding name, officials were more inclined to act on it than when it gave the same score to a teacher with a Dutch sounding name.
They called this selective adherence, and it has a consequence that goes well beyond the specific study. If adherence is selective, then an aggregate override rate tells you almost nothing. Ninety percent overrides does not mean ninety percent independent judgment. It could mean a reviewer who ignores the machine on the easy cases and defers on exactly the cases where deferring is most consequential. The pooled number and the dangerous behavior are perfectly compatible.
The number that would actually mean something is the override rate conditional on the machine being wrong. How often did the person catch the error? As far as I can find, almost nobody computes it, even though every organization running one of these workflows already holds the data in its logs.
The third failure is that the system moves underneath the reviewer.
This is the one the expo conversations pointed at without anyone quite saying it.
A person can only exercise judgment over a system whose behavior they roughly understand. But these systems do not hold still. The vendor ships a new model version. Someone edits a prompt. A criteria set updates. A retrieval corpus refreshes. None of those events is a new adoption decision, and none of them typically triggers a notification to the person whose name goes on the determination.
In United States healthcare the evidence that this causes real harm is already on the public record. When the Office of Inspector General examined Medicare Advantage prior authorization denials, it found that only thirteen percent met Medicare coverage rules, and among the causes it named systems that were not correctly programmed or updated. In October 2025 California regulators fined Cigna five hundred thousand dollars, partly because the company was running a review process different from the one it had filed with the regulator. The regulator noticed the drift. The plan did not.
Both of those are failures of the same kind. Something changed underneath a process that had been approved, and the approval stayed in place while the thing it approved quietly became a different thing.
The fourth failure is that nobody checks afterward.
Follow the third failure to its conclusion and you arrive somewhere genuinely empty. Requirements in this space oblige organizations to form committees, write policies, conduct annual reviews and maintain monitoring clauses. I have read a great deal of this material. I have found very little that obliges anyone to test, after a model changes or a criteria set updates, whether the decisions coming out are still the decisions the organization intended.
Even in medical devices, where the regulatory apparatus is strongest, the gap shows. The FDA’s final guidance on predetermined change control plans, published in August 2025, asks manufacturers to specify the methodology by which they will verify modifications. Of 794 AI enabled devices authorized between 2023 and 2025, forty three carried such a plan at all.
So the expo executives were right that their systems change. They were wrong, or at least unverified, in calling it correction. Correction requires a reference point outside the system. Retraining against your own objective and rescoring your own history is not a reference point. It is a mirror.
Debate
There is a strong case against everything I have just written.
Perhaps self-correction is real in some settings. In narrow domains with fast, unambiguous ground truth, machine learning systems genuinely do improve from feedback. A fraud model that learns from confirmed chargebacks has an external referee. A recommendation system that learns from purchases has one too. The label arrives, it was not generated by the model, and the loop closes honestly. So the claim I heard at the expo is not always false. It is true exactly where an outside signal exists and false where it does not, and the executives describing it rarely distinguished between the two.
Perhaps the meta-analysis does not apply. Vaccaro and colleagues studied experiments. Experimental subjects are often not domain experts, tasks are often artificial, and the stakes are low. Practitioners in real settings, with training and accountability and consequences, may behave better than undergraduates in a lab. That is a fair objection and I take it seriously. I would only note that the field studies we do have, including the mammography work and the prior authorization record, point the same direction rather than the other one.
Perhaps human oversight is not meant to improve accuracy at all. This is the most interesting version of the objection. On this reading, the human in the loop exists to provide a locus of responsibility, a route of appeal, and a check on the categorically bad outcome rather than the marginal one. Accuracy was never the point. Someone has to be answerable, and a machine cannot be.
I think this argument is partly right and it deserves to be said out loud more often than it is. But it comes with a price that its defenders rarely pay. If human review exists to hold responsibility rather than to improve decisions, then organizations should stop citing it as evidence that their decisions are sound, and regulators should stop accepting it as such. You cannot use the same control to claim both that a person catches the errors and that the person is there for accountability reasons unrelated to accuracy.
And perhaps the fix really is better design. Cognitive forcing functions, confidence display, deliberate friction, structured disagreement: researchers have tested all of these and some of them help. I hope that line of work continues. But every one of those interventions has to be verified in the specific deployment, which returns us to the same place. Somebody has to check whether the loop is doing anything, and that checking is what I cannot find evidence of anyone doing.
Reframe
Now the question I actually want to ask, which has been sitting underneath all of this.
If the loop is hollow, who is accountable?
The formal answer is clear and I do not dispute it. The organization made the decision. A person signed it. Law and regulation treat it that way and should. Nothing about deploying software transfers responsibility to the software.
The practical answer is where it comes apart, and the gap is not where people usually look for it.
The gap is not that nobody is responsible. The gap is that the responsible person has lost the capacity to exercise the responsibility they still carry.
Consider what that reviewer actually has. They have a recommendation from a system they did not build. The system was tuned by a vendor they have never met, using data they have never seen, against an objective nobody wrote down for them. The vendor changed the model six weeks ago and had no obligation to tell them. The reviewer has ninety seconds per case and a queue that does not shrink. The interface shows them a conclusion and not the basis for it. And their name goes on the outcome.
Responsibility without capacity is not accountability. It is exposure.
And it distributes strangely. The engineer who built the system is not accountable for the decision. The vendor who changed the model is not accountable for the determination. The executive who approved the deployment approved a version that no longer exists. The committee that met quarterly reviewed a policy rather than a behavior. Every one of those people can point, accurately, at someone else, and the only person who cannot point anywhere is the one with the least ability to have prevented the outcome.
This is what I think the expo executives were really telling me, without meaning to. When they said the system corrects itself, they were describing a world in which nobody has to hold the uncomfortable question. If the machine is self correcting, then the human does not need to catch it, and if the human does not need to catch it, then the loop can be thin without anyone feeling the thinness.
Self correction, told that way, is not a technical claim at all. It is a way of not having to ask who is watching.
Descent to Measurement
So what would it take to know?
Not to fix it. Just to know whether the loop in a given organization is doing anything.
Start with the smallest question. When a person reviews a machine recommendation, what do they do? Accept it as it stands, change it, or reject it and start again. Every system I have worked with already records enough to answer this, and very few organizations look.
Then the harder question, and it is the one that matters. Of the times the person changed the output, how often was the machine actually wrong? That single ratio separates a reviewer who is catching errors from a reviewer who is redoing good work, and it separates both of them from a reviewer who is approving errors. Three completely different situations with three completely different responses, and an aggregate override rate cannot tell them apart.
Then the direction of travel. Is that ratio moving? A reviewer who overrides less over time is either learning where the machine is reliable, which is good, or learning to stop looking, which is not. A snapshot cannot distinguish those. A trajectory can.
Then the thing that almost nobody tracks. When did the system last change, and what happened to the ratio afterward? If a model version ships on a Tuesday and the override rate drops on Wednesday and stays down, something happened and nobody in the organization currently has a way to notice.
None of this requires new instrumentation in most places. The delegation, the recommendation, the human action, the final output and the downstream result are already in the logs. What is missing is not data. It is anyone asking the logs a question about the person rather than about the model.
Measuring the human is uncomfortable territory and it should be. Nobody should build a leaderboard ranking employees on how well they defer to software, and any organization that tries will deserve the works council that stops them. The useful version points the other way: it tells the individual where their own judgment is adding value and where it is not, and it tells the organization which task types are genuinely suited to the machine and which are not. Those are different products, and the difference is who sees the name.
Close
I do not have this settled. That is my honest position.
What I have is a claim I heard repeatedly from people building serious systems, a body of evidence that does not support it, and a question underneath that I have not heard anyone answer well. If the loop is thinner than we say, then the person carrying the accountability is carrying it without the means to discharge it, and we have built a great deal on the assumption that they are not.
I am spending the next few weeks talking to people who run these workflows rather than people who sell them. Engineering leads, support leads, clinical informatics, operations, risk. If you have watched an AI output get rewritten and wondered why, or approved something you were not sure about, or discovered that a system changed underneath you, I would like twenty minutes of your time.
If a call is too much, the questions below take five minutes in writing and are just as useful to me.
Five minutes, if you have them
What AI system does your team rely on most, and what decision does it feed into?
Think of the last time someone on your team rewrote or rejected an AI output. Do you know why they did it?
Does anyone in your organization know how often the reviewer is right to disagree with the system? If so, who, and how do they know?
When the system last changed, whether a model version, a prompt, a rule set or a data source, how did the people reviewing its output find out?
After that change, did anyone check whether the decisions coming out were still the decisions you intended? Who, and what did they look at?
If an AI assisted decision at your organization turned out to be wrong, whose name is on it, and did that person have a realistic way to catch it?
What would you need to see to believe your human review step is actually working?
Reply to this email, or write to me directly. I will share what I learn across all the conversations, whether or not you take part.
References
Alberdi, E., Povyakalo, A., Strigini, L., & Ayton, P. (2004). Effects of incorrect computer-aided detection output on human decision-making in mammography. Academic Radiology, 11(8), 909–918.
Alon-Barkat, S., & Busuioc, M. (2023). Human–AI interactions in public sector decision making: “Automation bias” and “selective adherence” to algorithmic advice. Journal of Public Administration Research and Theory, 33(1), 153–169.
Agudo, U., Liberal, K. G., Arrese, M., & Matute, H. (2024). The impact of AI errors in a human-in-the-loop process. Cognitive Research: Principles and Implications, 9(1).
Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775–779.
Birhane, A., Steed, R., Ojewale, V., Vecchione, B., & Raji, I. D. (2024). AI auditing: The broken bus on the road to AI accountability. IEEE Conference on Secure and Trustworthy Machine Learning (SaTML).
Dzindolet, M. T., Peterson, S. A., Pomranky, R. A., Pierce, L. G., & Beck, H. P. (2003). The role of trust in automation reliance. International Journal of Human-Computer Studies, 58(6), 697–718.
Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1), 121–127.
Hoff, K. A., & Bashir, M. (2015). Trust in automation: Integrating empirical evidence on factors that influence trust. Human Factors, 57(3), 407–434.
Lyell, D., & Coiera, E. (2017). Automation bias and verification complexity: A systematic review. Journal of the American Medical Informatics Association, 24(2), 423–431.
Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410.
Skitka, L. J., Mosier, K. L., & Burdick, M. (1999). Does automation bias decision-making? International Journal of Human-Computer Studies, 51(5), 991–1006.
Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8, 2293–2303.
Original analysis. Written by Suneeta Modekurty, September 2026.


