Image generated using AI tools
I had conversations with two important people in governance and risk this week, and both left me thinking about the same question.
One of them has spent a career in healthcare, first on the side that pays for medicines and later on the side that regulates them. The other builds AI governance tools and has seen, at a scale most of us never will, how people actually use AI at work.
They came at the problem from very different directions.One looks at AI from the regulator’s side, where the question is: What needs to be true before this tool can affect a patient?The other looks at AI from the user’s side, where the question is: What do people actually do once the tool is in their hands?
I went into both conversations with a long list of questions. I came out with fewer answers than I expected, but with much better questions.
Three things stayed with me.
Accountability for AI is not one size.
Who is responsible for an AI tool depends on the organization using it, the agreement it has with the vendor, what the tool is being used for, and how much harm it could cause if it gets something wrong.
A ten-person clinic using a vendor’s note-taking assistant does not carry the same responsibilities as a large health system that builds its own model to prioritize emergency imaging.
For the small clinic, some of the protection comes from the contract it signs, the product it chooses, and the training it gives its staff. A health system that builds its own model takes on much more. It owns the testing. It owns the monitoring. It decides when the model is safe to use, and when it needs to be changed or switched off.
Both organizations are accountable. But they are not accountable for exactly the same things.
A governance system that treats them as if they are will fit neither one very well.
AI belongs to the whole organization, not to one department.
When an AI tool causes a problem in a hospital, the cause may not be purely technical or purely clinical.
The model may have performed exactly as it was designed to perform. The clinician may have used it carefully. And something can still go wrong in the space between the two.
Maybe a workflow was never redesigned. Maybe an alert appeared but nobody clearly owned the response. Maybe the vendor changed the system and the people using it were never told.
That space between technology and practice matters.
Organizations that handle AI well bring clinical and technical experts together before a tool is deployed, and they keep those people involved after deployment. When the AI comes from a vendor rather than being built internally, another party enters that relationship. The vendor now holds some of the knowledge about how the system works, while the healthcare organization holds the knowledge about where and how it is being used.
Responsibility is therefore shared, but it still has to be clear.
People use what helps them do the work.
Governance that sits in a policy document rarely changes what happens on a busy hospital floor or in a claims office at the end of a long day.
Guidance works better when it appears at the moment someone needs to make a decision. It needs to be inside the work, written in language people already understand, and designed so that doing the right thing is easier than ignoring it.
People do not always bypass governance because they are careless.
Sometimes they bypass it because the governance asks them to do more work without helping them do the work already in front of them.
When I put these three lessons next to each other, they lead to the same place.
There is a moment when an AI output stops being information on a screen and becomes something that happens to another person.
A note enters a medical chart.
A prescription goes to a pharmacy.
A coverage decision goes into a letter.
A message is sent to a patient.
That is the point where accountability becomes real. It is where clinical work, technology, governance and responsibility meet. It is where guidance either helps or gets ignored.
And it is where human judgment matters most.
The rest of this piece is about that moment, what several recent pieces of work tell us about it, and why I think healthcare may be getting something important right.
What the UK put on paper
On September 10, 2026, the National Commission into the Regulation of AI in Healthcare published its report.
The Medicines and Healthcare products Regulatory Agency created the Commission a year earlier, in September 2025, and asked it to advise the UK government on a new regulatory framework for AI in healthcare.
The final report is 119 pages long and contains 44 recommendations. Those recommendations came from a broad process that included a public call for evidence, discussions with patients and the public, sessions with groups whose voices are often missed, roundtables with healthcare professionals and industry, and specialist working groups.
The executive summary begins with a sentence I think every healthcare leader should read carefully.
AI should augment, rather than replace, healthcare professionals.
The Commission then explains what that can look like in practice. AI can take over some routine and administrative work, giving healthcare professionals more time for the things that remain deeply human: communication, compassion and shared decision-making.
The evidence collected by the Commission points in the same direction.
People are not simply rejecting AI in healthcare. But their support is not unconditional either.
They want to know how AI will be used safely, fairly and transparently. They want evidence that it actually helps. They want meaningful human oversight. And they want to know who is accountable when something goes wrong.
The Commission’s central argument is that AI regulation in healthcare needs to become more proportionate, lifecycle-based and system-wide.
Those can sound like regulatory words, but each one means something quite practical.
Proportionate means that the amount of scrutiny should match what the tool can do, the benefit it offers and the harm it could cause.
An AI tool that drafts a discharge summary for a nurse to review does not need exactly the same level of scrutiny as a system that decides which chest scans a radiologist should see first.
The consequences are different. The oversight should reflect that.
Lifecycle-based means that approving an AI system is not the end of the job.
It is the beginning.
AI products can change. Their performance can change when they are used in a new hospital or with a different patient population. Their behavior also depends on the data, workflows, people and organizations around them.
So governance cannot stop on the day a tool is approved.
It has to follow the tool through deployment, real-world use, monitoring, updates and whatever is learned along the way.
System-wide means regulators cannot carry the entire responsibility themselves.
Manufacturers have responsibilities. Healthcare organizations have responsibilities. Healthcare professionals have responsibilities. Regulators and policymakers have responsibilities.
Several of the Commission’s 44 recommendations make these responsibilities very concrete.
Recommendation 26 asks the MHRA to require manufacturers to state the conditions that need to be in place for their product to be used safely.
The Commission gives examples such as cybersecurity, user training and the healthcare organization’s own readiness to use AI.
Put simply, the manufacturer should not only say, “Here is our AI tool.”
It should also say, “Here is what needs to be true in your organization for this tool to be used safely.”
Recommendation 28 goes one step further. It asks for contracts between manufacturers and healthcare providers to clearly state who is responsible for each of those safeguards.
That matters because “someone should make sure this happens” is not the same as knowing who that someone is.
Recommendation 29 asks UK health departments to develop a readiness toolbox. The idea is to help a healthcare provider determine whether it is actually ready to deploy a particular AI product, whether it can provide the safeguards that product requires, and whether it can clearly describe the governance around its use.
Recommendation 33 focuses on the people using the technology. It asks healthcare providers to make sure staff are properly prepared, including training on the specific AI tools they use in practice.
Recommendation 34 asks for a strong culture of reporting what happens after deployment, including adverse outcomes and product malfunctions. That creates a way for organizations and regulators to keep learning how AI performs in real healthcare settings rather than relying only on what was known before launch.
Recommendation 35 focuses on patients. It asks healthcare organizations to be transparent about the use of AI in care, including the reasonable expectation that patients will be told about its use and, where possible or appropriate, have an opportunity to opt out.
Put these recommendations together and a clear picture starts to appear.
This is not a system in which AI makes decisions while people stand nearby and watch.
It is a system in which people are trained on the specific tools they use. The safeguards around those tools are written down. Responsibility for those safeguards is assigned. Problems are reported and learned from. And patients are told when AI is part of their care.
Human judgment is not something added at the end.
It is part of the structure holding the system together.
I am writing this from the United States, and this is a UK report. These are recommendations to the UK government, and a formal government response is still to come.
But the questions behind the recommendations are not uniquely British.
A hospital in Missouri eventually faces the same basic questions as a hospital in Manchester:
Which AI tools are we using?
Who has been trained to use them?
Who is watching how they perform?
And who decides when an AI output is allowed to become something that affects a patient?
How close the human stands depends on what is at stake
The word in the Commission’s report that I keep coming back to is proportionate.
It gives us a useful way out of a debate that is often framed too simply.
Should a human review everything an AI system produces?
Should AI ever be allowed to act without a human approving every action?
Ask the questions that way and we quickly end up arguing in absolutes.
Ask instead what happens when the AI is wrong, and the question becomes much more practical.
In healthcare, the default should remain human judgment. Clinicians decide and AI assists.
But that does not mean a person needs to stand equally close to every single AI output.
How close that person needs to be depends on what is at stake.
Consider four kinds of AI tools that a mid-sized health system might use today. These are examples to explain the principle, not a published standard.
At the lower end, imagine an agent that sends appointment reminders and answers questions about parking.
If it gives someone the wrong parking information, the patient may be frustrated or inconvenienced. That matters, but the likely harm is limited.
A person might therefore review the message templates before the system goes live, check a sample of messages every month, and review complaints.
Nobody needs to manually approve every parking message before it is sent.
Now consider an ambient scribe that listens during a medical visit and drafts the clinical note.
The stakes are higher.
An incorrect note can remain in a patient’s medical record for years. It can influence referrals, insurance decisions and future treatment.
Here, the clinician should read the note and sign it before it becomes part of the chart. The organization should also know where the audio is stored, where the draft is stored, and how that information is handled.
Now move higher again.
Imagine a triage system that changes the order of a radiologist’s worklist by moving scans with possible urgent findings toward the top.
If the system moves a serious case down instead of up, a patient may wait longer for important care.
The radiologist still reads every study. But that alone is not enough.
Someone also needs to monitor how the system performs across many cases. The organization needs to know whether it is missing cases it should catch and whether those failures differ across patient groups. It needs to keep watching because performance can change as patient populations, equipment or workflows change.
Finally, consider an AI system that recommends a medication dose or recommends denying coverage for a medical procedure.
Now the output directly changes what happens to a person’s body or whether that person can access care.
At this level, a qualified person should make the decision in each case. The reasons should be recorded. And the patient should have a way to question or challenge the decision.
Notice what changes as we move up this ladder.
It is not whether a human is involved.
A human is involved at every level.
What changes is where the human stands.
For one tool, a person may review a template before launch.
For another, a clinician signs every output.
For another, people monitor patterns across thousands of outputs.
And for the highest-risk decisions, a qualified person makes the decision case by case.
That is what proportionate oversight means to me.
It does not mean less oversight.
It means putting oversight where it can actually prevent harm.
This also brings me back to my first takeaway from those conversations.
The same AI product may carry different levels of risk in different organizations and different settings.
An ambient scribe used in a cosmetic dermatology practice and the same scribe used in an oncology clinic may create very different consequences because the notes are being used in different kinds of care and can lead to very different decisions.
Risk does not belong only to the tool.
Risk also belongs to the way the tool is used.
What one builder chose to keep human
On September 29, 2026, Meta announced Muse for Small Business, an extension of its Muse AI agent that connects with software already used by small businesses.
Reporting on the launch described connections with Facebook and Instagram business accounts and services such as Shopify, Stripe, QuickBooks, Canva and Slack.
The agent can use information from a business’s storefront, financial records and customer information to analyze sales, look for unusual expenses and prepare drafts.
Muse is not a healthcare product, and I am not presenting it as one.
What interests me is one design decision Meta made because it shows where a builder chose to draw the line between what an AI agent can do by itself and what still requires a person.
According to Meta, Muse does not publish, send or spend without the business owner’s approval.
When Muse carries out a transaction with a new merchant, a Meta spokesperson also told Retail Dive that it uses a single-use card number rather than collecting the owner’s password or payment information.
Look closely at where the human approval appears.
The agent is allowed to read.
It is allowed to analyze.
It is allowed to draft.
It can look across sales, advertising and accounts and come back with recommendations or a plan.
But when the action leaves the system and reaches the outside world, a person steps in.
A public post requires approval.
A message that reaches a customer requires approval.
Money moving out of the business requires approval.
Those are moments of consequence.
And those are the moments Meta chose to keep human.
That is striking because it resembles the line the UK Commission is drawing in healthcare, even though the two are working in completely different settings and for completely different reasons.
The agent prepares.
The person decides when the action is about to affect someone or something outside the system.
When a technology builder and a healthcare regulator arrive at a similar boundary from opposite directions, I think the boundary is worth paying attention to.
What the agents show about themselves
Some of the strongest evidence I have seen this year for keeping human judgment close to important decisions comes from research designed to test AI agents themselves.
In April 2026, Jeremy Li and Andrew Ho at OpenAI published GeneBench, a benchmark designed to test AI agents on realistic, multi-stage data analysis in genetics and quantitative biology.
Many earlier biology benchmarks asked whether a model could recall facts or perform a familiar analytical task.
GeneBench asks something harder.
It asks an agent to do more of the work that a computational biologist actually does: take messy data, inspect it, choose an analytical method, run the analysis and reach a scientific conclusion.
The first version contains 103 problems across ten domains. Many of those problems involve decision points that resemble the choices scientists make in genetics-backed drug discovery and translational research.
In June 2026, the authors released GeneBench-Pro, an expanded version with 129 problems across genomics, quantitative biology and translational biomedicine.
The finding I find most important is not the overall score.
It is a pattern in how the agents reason.
The authors found that models can complete large parts of these workflows, but they repeatedly show a gap between noticing something and acting on what they noticed.
An agent may correctly detect a warning during a diagnostic check. It may notice a strange pattern in the residuals or a quality problem in a sample.
But then it may fail to carry the meaning of that warning into the next decision.
It noticed the problem.
It did not change what it did because of the problem.
That can lead the agent to choose the wrong statistical method or continue down an analytical path that looked reasonable at the beginning but is no longer appropriate given what the data revealed.
I find that result both humbling and clarifying.
It is humbling because these are capable systems, and the people building them are serious researchers who chose to measure and publish where the systems still fall short.
But it is also clarifying because it helps us describe what human judgment actually contributes.
The difference is not simply that humans notice things and AI does not.
AI can notice.
The harder step is connecting what was noticed to what should happen next.
Now translate that idea into healthcare.
An AI system reviewing a patient’s record may notice that a laboratory value is outside the normal range.
It may also notice that the patient recently started a new medication.
But the important question is whether it connects those two facts and understands that the abnormal lab result should change the dose it is about to recommend.
A pharmacist or physician may see those same two pieces of information and immediately understand that one changes the meaning of the other.
That connection is judgment.
And it is the kind of step GeneBench shows that agents do not yet handle reliably.
GeneBench studies scientific analysis, not bedside medicine, and I do not want to claim that it proves something it did not test.
But the gap it identifies—the gap between noticing evidence and changing the next action because of that evidence—is not unique to genetics.
It appears whenever a decision requires someone to carry information from one step into the next.
Much of medicine works exactly that way.
What patients are asking for
In June 2026, the Centre of Excellence for Regulatory Science and Innovation in AI and Digital Health, or CERSI-AI, held three workshops on AI and professional liability in health and care.
The first brought together patients and members of the public.
The second included medical defense organizations, insurers, professional bodies, unions, law firms and academics.
The third included professional regulators, health service regulators and others responsible for setting standards of care.
CERSI-AI published the findings in September 2026.
The report cites public polling showing that between 75 and 79 percent of patients believe they should be told in advance when AI is being used to support diagnosis and triage.
Patients who participated in the first workshop also said they wanted to understand how AI was being used in their care and how they could raise a concern.
The workshops proposed that healthcare providers publish information about which AI tools they use and where they use them.
That information would also explain what patients can expect, how the tools are monitored and what happens when something goes wrong.
Again, what patients are asking for is not simply the removal of AI.
They are asking for someone to remain accountable for what happens.
They want a person who understands how the tool was used, can explain its role and can stand behind the final decision.
When a patient asks, “Who decided this?”, the answer they are looking for is a person, not the name of a model.
But there is another side to this.
The same CERSI-AI report records concerns from clinicians.
Clinicians worry that they may be held accountable for AI tools whose performance they cannot personally verify and that were selected through procurement processes in which they had little or no involvement.
That is a real problem.
The workshop participants treated it as a foreseeable system-level risk rather than something that should simply be pushed onto the individual clinician.
Responsibility for AI cannot end with the person standing in front of the patient if that person had no meaningful role in selecting, testing, monitoring or understanding the tool.
I will return to the clinician’s side of this question in the next issue because I think it deserves an article of its own.
What human judgment actually is
“Human judgment” is an easy phrase to use.
It can also become meaningless if we never explain what we mean by it.
So here is what I mean when I use the term in this article.
Judgment is noticing what matters in front of you.
A clinician looks at a patient and realizes the patient appears much sicker than the numbers on the screen suggest.
Judgment is connecting what you noticed to what you are about to do.
A pharmacist sees a new kidney result and understands that it changes the dose that should be given.
This is the kind of step GeneBench found agents can miss.
Judgment is deciding, including deciding not to follow the expected path.
A radiologist reads a scan that the triage system ranked as routine and decides to call the ward anyway.
Judgment is owning the decision afterward.
A physician signs the note, explains the plan to the patient and remains accountable for the decision if something goes wrong.
The authority to do these things is what I call cognitive authority.
It is the ability and the standing to look at an AI output and say:
Yes.
No.
Not yet.
And it includes having enough knowledge and skill to know which answer is appropriate.
In a well-designed AI system, AI can make parts of this process faster and richer.
It can surface information a person might otherwise have to spend much longer finding. It can bring patterns forward. It can organize evidence. It can show possibilities and help a person see more than they could easily see alone.
But the final steps still matter.
Someone has to decide what the information means in this particular situation.
Someone has to decide what happens next.
And if we expect a person to carry that responsibility, the system has to give that person enough time, information and skill to actually exercise judgment.
We cannot keep a human “in the loop” in name while designing the work so that the human has no realistic opportunity to question the machine.
Keeping cognitive authority is also not automatic.
Skills that are not used can weaken.
If a person accepts the AI recommendation every time, questioning it can slowly stop being part of the work.
That does not mean we should slow AI down or refuse to automate useful work.
It means we should design the work carefully.
People need to remain engaged with the decisions for which they are accountable.
And they need training on the specific tools they are expected to judge.
That is also why Recommendation 33 in the UK Commission’s report matters. General AI literacy is useful, but it is not enough. People need to understand the tools they actually use in their work.
What I think, and what I will keep writing about
After these conversations and after reading this work, here is where I have landed.
I think healthcare may be ahead of many other industries on this question, and not simply because healthcare is cautious.
Medicine has always worked with tools that can do particular things better than humans can.
A laboratory analyzer measures potassium more precisely than a clinician ever could.
An MRI scanner lets us see what the human eye cannot see on its own.
Healthcare learned long ago that using a powerful instrument does not require handing the entire decision to the instrument.
We trust the instrument to do what it does well.
Then a trained person interprets what it means for this patient.
AI agents are far more capable and flexible than traditional medical instruments, but I think the basic principle still holds.
The tool informs. The person decides.
I also think the right amount of human involvement should be determined by risk rather than ideology.
In general, a human remains involved.
But how close that person needs to stand depends on what can happen to a patient when the AI is wrong.
Appointment reminders and scheduling may run with relatively light review.
Clinical notes need a clinician’s signature.
Triage systems need ongoing monitoring.
Medication doses and coverage decisions need qualified people making decisions about individual cases.
The human does not disappear as automation increases.
The human moves to the places where judgment carries the most consequence.
Regulators in the UK are beginning to put versions of that principle into policy. Technology builders such as Meta are making similar choices in very different products.
I also think one of the most important gaps we should watch is not simply the gap between human intelligence and artificial intelligence.
It is the gap between noticing and acting.
AI agents are becoming very good at finding things.
They can surface evidence, identify patterns, retrieve information and point out abnormalities.
But noticing something is not the same as understanding what that thing should change.
Connecting new information to the next decision, in the context of one particular patient, remains a critical part of human judgment.
Finally, I think organizations have to protect that judgment deliberately.
Human judgment will not remain strong simply because a policy says a human is accountable.
It has to be supported through training on specific tools.
It has to be supported by workflows that keep people engaged with the decisions they are expected to own.
Each safeguard needs a clear owner.
Organizations need to know what changes when vendors update their products.
And when something goes wrong, people need to be able to report it honestly so that the organization can learn from it.
That is what meaningful human involvement looks like to me.
Not a person placed at the end of an automated process simply to click approve.
A person who understands what the AI did, understands what is at stake, has the authority to disagree, and is close enough to the consequence to intervene when intervention matters.
In the coming issues, I want to keep writing about healthcare and about how humans and AI agents will share this work.
I want to look more closely at the clinician’s side of accountability.
I want to look at how an organization decides exactly where a human needs to stand for each AI tool.
And I want to explore what it takes to keep human judgment sharp when the AI is right most of the time—because that may be one of the hardest problems of all.
I will continue grounding these pieces in published research and documented events so that the arguments are not simply opinions about where AI might be going.
They should be ideas you can check, question and take into your own conversations.
Four questions to bring to your next AI conversation
If you lead a healthcare organization, work inside one, or advise one, there are four questions I would ask about every AI tool that is already being used or is about to go live.
Where does this tool’s output reach a patient?
Find the exact moment.
Is it when a note enters the medical record?
When a message is sent?
When a medication dose is suggested?
When a coverage letter goes out?
Do not stop at saying that the AI is “used in the workflow.”
Find the point where what the AI produces becomes something that can affect a person.
That is the point of consequence.
And that is where you need to know who is standing there.
How close does that person need to stand, and why?
Not every AI output needs the same kind of human review.
For one tool, reviewing templates before deployment may be enough.
For another, someone may need to monitor performance across thousands of cases.
For another, a qualified person may need to read and sign every output.
And for some decisions, a person may need to decide every case individually.
The reason should be clear: What happens to the patient if this output is wrong?
That answer should determine how close the human stands.
Who has been trained on this specific tool?
General AI education matters.
But knowing broadly how AI works is different from understanding the AI system you use every day.
People need to know what their particular tool can do, what it cannot do, where it tends to fail, what information it uses and what they are expected to do when something looks wrong.
Know who received that training.
Know when they received it.
And know whether the training changes when the tool changes.
When the vendor updates the tool, who finds out?
An AI system that was tested and considered safe when you approved it may not remain exactly the same system.
Models change.
Features change.
Data sources change.
Interfaces change.
Behavior can change.
Someone inside the organization needs to know when an important update happens, understand whether it changes the assumptions under which the tool was approved, and decide what needs to be checked before that change reaches patients.
None of these questions is meant to slow AI down.
They are meant to make sure we know what happens when AI leaves the screen and enters someone’s life.
Because at that moment, the most important question may not be how capable the AI is.
It may be whether a person who understands both the tool and the consequence is still close enough to make the call.
Sources
National Commission into the Regulation of AI in Healthcare, Full National Commission Report, Medicines and Healthcare products Regulatory Agency, September 2026. Executive summary; Recommendations 26, 28, 29, 33, 34 and 35 in Annex A.
CERSI-AI, AI and professional liability in health and care: Opportunities for action, September 2026.
Jeremy Li and Andrew Ho, “GeneBench: Assessing AI Agents for Multi-Stage Inference Problems in Genomics and Quantitative Biology,” bioRxiv, April 2026.
Jeremy Li and Andrew Ho, “GeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicine,” bioRxiv, June 2026.
Help Net Security, “Meta gives small businesses an AI agent that knows their work,” September 29, 2026.
Retail Dive, “Meta debuts AI agent for small businesses,” September 2026.


