
On May 22, 1918, the Madrid newspaper “ABC” reported that a strange illness was spreading through the city. People were falling sick in large numbers, doctors did not yet understand what they were dealing with, and the newspaper reported what it knew.
There should be nothing remarkable about this. A newspaper found an outbreak in its city and wrote about it. What made it remarkable was what was happening elsewhere.
Europe was still at war. Britain had the Defence of the Realm Act, France and Germany operated under wartime censorship, and the United States had its own restrictions on information considered damaging to morale. The reasoning was understandable. Armies were still fighting, and governments needed civilians to keep producing, soldiers to keep fighting, and populations to believe that the war remained winnable. Reports of a disease tearing through military camps and killing young soldiers were not exactly helpful to morale. So information about the epidemic was restricted, softened or delayed.
Spain had a different problem. It was neutral. There was no Spanish war effort whose morale needed protecting, so its newspapers could report the outbreak much more openly, and when King Alfonso XIII became seriously ill, the Spanish press reported that too.
The world suddenly had a great deal of information about influenza in Spain. It had much less information about influenza elsewhere.
And a strange thing happened. The country that talked about the disease became associated with the disease.
The pandemic would eventually kill tens of millions of people around the world. Its precise origins remain debated. It certainly did not respect national borders, and Spain was not uniquely responsible for it. Yet more than a century later, we still know it as the Spanish flu.
There are several reasons historical names take hold, and the history is more complicated than a single newspaper report. But one feature of this story has stayed with me. The place that was easier to see became easier to blame.
Spain reported what was happening. Other countries reported under very different constraints. Put those reports beside one another without understanding how they were produced, and you could easily draw the wrong conclusion. Not because Spain’s numbers were false, but because the numbers were unequally visible.
Organizations are building AI measurement systems right now that may encounter a modern version of the same problem.
The clean company and the honest company
Forget AI for a minute. Imagine you manage two teams, and every Friday you ask both managers to send you a list of unresolved problems.
The first manager is obsessive. She talks to everyone. She checks the support queue. She asks engineers about workarounds. She looks at customer complaints. She finds the spreadsheet somebody has been quietly maintaining because the official system doesn’t capture a certain failure.
Friday afternoon, her report lands in your inbox. Thirty-seven problems.
The second manager asks the people sitting nearest to him whether anything major happened. Nobody mentions much. His report arrives. Four problems.
If all you have are those two numbers, which team looks healthier? The second one. Which manager is doing the better job of finding problems? Probably the first.
Now imagine bonuses depend on the number. You have just created an incentive for the first manager to become a little more like the second.
This is an old organizational problem. We encounter it in safety reporting, cybersecurity, quality control, healthcare, employee complaints and incident management. Sometimes a rise in reported problems means conditions are deteriorating, and sometimes it means detection improved. Sometimes a fall means things got better, and sometimes it means people stopped looking. The number alone cannot tell you which world you are in.
That is why measurement isn’t simply about counting. It is about understanding what had the opportunity to be counted.
What was the number actually measuring?
Suppose you had been sitting in London in late 1918 with a table of reported influenza cases by country, and someone points to the columns and asks what the table measures.
The obvious answer is influenza. But the numbers would also have reflected something else, which is the conditions under which influenza could become visible. Disease prevalence mattered. So did surveillance, and reporting, and censorship, and the willingness and ability of institutions to say publicly what they knew.
The number was not floating independently in the world. It had been produced by a system, and that system determined what could enter the numerator in the first place.
Which brings us to a deceptively simple idea. A number is inseparable from the population that had a chance to become part of it.
Say twenty AI systems violated policy last month. Twenty out of what? Twenty out of twenty-five? Twenty out of two thousand five hundred? Twenty out of every AI system operating in the company, or twenty out of every AI system the governance team knew existed?
Those are not different interpretations of the same measurement. They are different measurements.
And yet in executive reporting, denominators have a habit of disappearing. The numerator survives. It gets a dashboard. It gets a green arrow. It gets compared with last quarter. It gets quoted in a board meeting. Eventually somebody says that only three percent of our AI applications have critical governance findings.
Three percent sounds precise. But before I know what to think about three percent, I need another number. Three percent of what? And then another question. Who decided what belonged in the what?
That second question is where things become uncomfortable.
A very good answer to one AI governance problem
Microsoft recently published an architecture piece called *From Policy to Proof: Governing AI to Scale Human Ambition and Machine Intelligence*, and there is an important idea at the centre of it. A policy document is not evidence that governance is happening.
A company can write an excellent AI policy. It can assign owners, establish committees, document approval processes and require employees to follow rules. None of that, by itself, proves what happened when an AI system actually ran.
Runtime governance tries to close that gap. Instead of leaving governance in documents, controls move closer to production. Rules become enforceable baselines. Traffic to models, agents, APIs and tools can pass through gateways that enforce authentication, authorization, quotas and other policies. Agents can be checked at different stages of their lifecycle. Models can be evaluated before release and monitored after deployment. Observability records what happened, and audit systems turn those events into evidence.
This is a substantial improvement over the world in which somebody writes a policy in January and somebody else completes a questionnaire in December saying the policy exists. A production system can now say: this request happened, this control was applied, this action was blocked, this exception occurred, this person approved it, this was the output.
That is real evidence, and the underlying principle is exactly right. Your AI policy is not governance merely because it exists. Production should be able to show that the policy operated.
But solving one measurement problem often exposes the next one. And the next one begins before the gateway. It begins with the list.
Before you govern AI, you have to decide what counts as your AI
The recommended governance sequence begins sensibly enough. First, establish scope, and inventory the AI applications, Copilots, agents, APIs, tools, MCP servers and custom workloads operating across the organization. Then classify them. Then apply controls. Then monitor. Then generate evidence.
The engineering can be excellent all the way down. But read the first step again. Inventory the AI estate. Everything that happens later depends on what entered that inventory.
A gateway can enforce policy on traffic that reaches the gateway. An evaluation system can evaluate models it knows about. An audit system can prove that controls fired on workloads within its boundary. Observability can provide extraordinarily detailed evidence about the systems it observes. But none of those systems can independently prove that the inventory represents the entire AI estate.
This isn’t a criticism of the technology. A metal detector at an airport cannot tell you how many people entered through a door that bypassed the metal detector. That doesn’t make the metal detector defective. It means coverage is a different measurement problem from detection, and enterprise AI is going to have to learn that distinction.
Imagine these two companies
Company A takes AI inventory seriously.
The team starts with the sanctioned systems everyone knows about. Then somebody asks an irritating question: what are we missing?
So they keep looking.
They talk to business units.
They inspect old pilots.
They find the claims experiment that was supposed to end six months ago but never really did.
They discover two MCP servers an engineering team stood up.
They find employees paying for AI tools on corporate cards.
They find an agent built inside an existing enterprise licence that central IT did not realize had become part of an important workflow.
They find a team moving customer information through a workflow nobody designed centrally, because the team was simply trying to get its job done.
Six weeks later, Company A has a rather ugly inventory.
Now governance gets applied. There are exceptions. There are unapproved tools. There are controls that don’t fire. There are workloads that need remediation. There are gaps. There are uncomfortable conversations. The audit report is full of them.
Now Company B.
Company B also performs an inventory. The platform team lists the AI systems the platform team deployed. Everything is sanctioned. Everything is registered. Everything is already behind the appropriate infrastructure. Controls are applied. Evaluations run. The audit trail looks clean.
Put the two reports in front of a board. Which company looks better?
Company B.
Put them in front of an insurer.
Company B.
A large customer doing vendor diligence?
Probably Company B.
An acquiring company?
Again, Company B.
But there is one thing we don’t know. Which company is actually better governed?
Company A found more problems because Company A looked harder. Company B may genuinely have a cleaner AI estate, or Company B may simply have measured a smaller part of it. The governance reports alone cannot distinguish those possibilities.
The company that looks worse may be the company that looked harder.
Spain reported more visibly. Company A discovered more aggressively. Different century, different stakes, different mechanism, and the same measurement warning. Before comparing the findings, compare the opportunity to find them.
Nobody needs to cheat for this to go wrong
It would be easy to turn this into a story about companies manipulating their AI inventories. That would make for a dramatic article, and it would miss the more interesting problem. Most organizations don’t need to manipulate anything, because the inventory can become incomplete through completely ordinary behaviour.
Someone has three weeks to finish it. They start with procurement records. Another person starts with the CMDB. Another asks department heads to submit their AI applications. The platform team exports everything registered in its environment. Security contributes the services it knows about. Legal contributes the vendors whose contracts mention AI. Everyone works hard. Nobody lies. And still, the resulting inventory may be incomplete.
Because AI use does not necessarily begin with an AI deployment. An employee buys a subscription. A vendor quietly adds AI to a product the company has used for five years. A developer calls an external model API. A department builds an agent inside software already approved for another purpose. An old pilot keeps running. A feature that used to be deterministic becomes AI-enabled after a vendor update. A workflow moves from experimentation to dependency without anyone formally declaring that transition.
The organization’s AI estate changes faster than its inventory process.
Now add incentives. Suppose companies with fewer governance findings receive easier approvals, or lower insurance scrutiny, or faster customer procurement, or better board reporting. Nobody has to issue an instruction saying please find less AI. Organizations learn what their measurement systems reward. The exhaustive inventory creates work. The narrow inventory creates green dashboards. Over time, that matters.
There is a principle from measurement theory and organizational behaviour hiding here. If finding more of the truth makes your score worse, eventually the measurement system starts discouraging discovery. Not because everyone becomes dishonest, but because people respond to incentives.
Why sophisticated buyers may still miss it
You might think a board, insurer, regulator or enterprise procurement team would notice. Sometimes they will. But runtime governance creates an unusual psychological problem, which is that the evidence looks excellent. And much of it is excellent.
Machine-generated logs are persuasive. They have timestamps. They identify systems. They show controls firing. They show exceptions. They can demonstrate that a particular request passed through a particular policy at a particular moment. Compare that with the artifacts organizations have historically used to demonstrate governance: a policy PDF, a spreadsheet, a questionnaire, a screenshot, an annual certification, an employee saying yes, we have a process for that.
Runtime evidence is obviously better. And sometimes being obviously better than the alternative prevents us from asking what comes next.
Imagine receiving eighty thousand rows of enforcement telemetry. The volume itself feels like coverage. The specificity feels like rigour. The automation feels objective. But there is a question that eighty thousand rows cannot answer merely by becoming eight hundred thousand rows. What proportion of the organization’s actual AI activity had the opportunity to generate one of these rows?
That is a coverage question, and a system operating inside its own boundary cannot necessarily answer it. A gateway knows what passed through the gateway. It does not, by itself, know what bypassed it. An inventory knows what was inventoried. It does not know what nobody thought to inventory. An audit can prove that a control fired. It cannot automatically prove that every relevant system was subject to that control.
These are different claims, and we should stop treating them as one.
There are really two questions
When somebody tells me their AI systems are governed, I increasingly hear two separate questions hiding inside the sentence.
The first is whether the systems inside the governance boundary are actually governed. This is where runtime controls, evaluations, observability and audit evidence can be extremely powerful.
The second is how much of the actual AI estate is inside that boundary. That requires a different kind of evidence.
You can perform the first beautifully and still not know the answer to the second. The distinction seems almost embarrassingly obvious once written down, but many measurement failures are like that. The denominator becomes visible only after someone asks for it.
What eventually corrected the picture of 1918?
Historians did not reconstruct the pandemic simply by finding better newspaper articles. They looked at other evidence: military medical records, hospital admissions, municipal records, burial registers, mortality statistics, and, critically, excess mortality, meaning how many more people died than would normally have been expected.
These sources had different purposes. A burial register was not designed to settle an argument about international influenza prevalence, and that is part of why it was useful. It gave researchers another view of reality. No single source was perfect, but multiple sources made it harder for the visibility of one reporting system to masquerade as the underlying phenomenon.
That idea transfers well to enterprise measurement. When the primary measurement system has a blind spot, don’t merely ask it to describe the blind spot better. Look for another source.
For AI, some of those sources may already exist. Identity systems know which accounts authenticate to services. Expense systems know what the company is paying for. Vendor records know which services have contracts. Network telemetry may reveal communication with AI endpoints. Browser and endpoint systems may reveal another part of the picture. Data security systems may see yet another. Business teams know which workflows have quietly become dependent on AI.
None of these sources alone gives you the AI estate. They all have blind spots. But that is precisely the point, because independent imperfections can be more useful than one beautifully instrumented boundary.
If the inventory says a hundred and forty AI systems exist, while procurement suggests a hundred and eighty-five AI-related vendors and identity telemetry shows employees authenticating to another set of services, the discrepancy is information. It doesn’t immediately tell you which number is correct. It tells you where to look. And that may be more valuable than another decimal place on the governance score.
The denominator needs evidence too
We spend enormous energy proving numerators. How many systems passed evaluation, how many failed, how many policy violations occurred, how many AI applications are approved, how many incidents were detected, how many controls fired.
But the denominator often arrives almost casually. Across our AI estate. Of all AI applications. Enterprise-wide. All production AI.
Those phrases sound like scope descriptions. They are actually claims. All AI applications is itself something that needs evidence.
That may be one of the most important measurement principles for enterprise AI. Coverage should be measured, not assumed.
If I tell you that ninety-seven percent of our AI systems meet a governance requirement, there are two measurements hiding inside that statement. The first is that of the systems we evaluated, ninety-seven percent met the requirement. The second is that the systems we evaluated represent the AI estate we claim they represent.
The first may be generated automatically. The second requires independent justification. Without it, ninety-seven percent can be completely accurate and still create false confidence. The number can be true and the interpretation can still be wrong.
The stranger test
There is a simple way to tell whether a measurement is strong. Give it to someone who was not in the room when it was created.
Your internal team knows what AI estate means. They know which departments participated in the inventory. They know that one subsidiary was excluded because its systems are being migrated. They know the sales team submitted its inventory late. They know the customer-service pilot wasn’t included because technically it is still considered experimental. They know all the footnotes.
Then the number leaves the room.
It reaches the board. Ninety-four percent governed.
It reaches a customer. Ninety-four percent governed.
It reaches an insurer. Ninety-four percent governed.
It reaches an acquirer. Ninety-four percent governed.
The context gets stripped away while the precision remains. That is when the stranger asks the question the internal organization stopped asking, because everybody already knew the answer.
Ninety-four percent of what?
And then: how do you know that’s all of it?
This is going to matter more, not less. More customers are asking vendors about AI. Boards are asking management what is being deployed. Insurers are trying to understand emerging exposure. Acquirers are going to want to know what AI systems, agents and dependencies they are inheriting. Regulators will ask their own versions of these questions.
The person receiving your measurement increasingly will not be the person who designed it. That changes what a good measurement has to survive. It has to survive the stranger.
A clean dashboard should make you curious
There is a cultural implication here too. Organizations naturally celebrate green dashboards, and we should probably be a little more suspicious of them. Not because green is bad, but because the first period of serious measurement often makes an organization look worse.
A company installs better cybersecurity monitoring and suddenly detects more incidents. Did security deteriorate overnight? Maybe. Or perhaps visibility improved.
A hospital creates a stronger reporting culture and receives more safety reports. Did the hospital become less safe? Not necessarily. People may finally feel able to report what was already happening.
An AI governance team expands discovery and finds twice as many unsanctioned systems. Did AI governance fail? Possibly the opposite. The organization may finally know what it is governing.
This creates a difficult executive problem. How do you reward discovery without rewarding failure? If every newly discovered problem makes a leader’s score worse, you have built a system that punishes visibility, and eventually people learn. Don’t look too hard. Don’t classify ambiguous cases. Don’t expand scope until next quarter. Don’t turn the pilot into an official production system yet. Don’t add the strange edge case to the inventory until we understand it.
The measurement itself teaches the behaviour.
That is why a mature measurement system should be able to distinguish at least two things: the state of what we can currently see, and our confidence that we are seeing what matters. Those are not the same score, and they should not be treated as one.
The company with more findings may be healthier
Suppose Company A discovers forty-seven governance gaps and Company B discovers six. Without coverage information, I cannot tell you which company has the stronger governance programme. Forty-seven findings could mean poor governance, and they could also mean excellent discovery. Six findings could mean excellent governance, and they could also mean poor discovery.
The findings are not useless. Far from it. They are simply incomplete without information about how comprehensively the organization looked.
That means an AI measurement system should not merely ask how many problems we found. It should also ask how hard we looked, where we looked, what independent sources corroborate the boundary, what might remain outside it, and how our coverage has changed since the last measurement.
Now an increase in findings can be interpreted properly. Perhaps the organization genuinely deteriorated. Perhaps its AI footprint expanded. Perhaps the measurement became more sensitive. Perhaps discovery coverage improved. Perhaps all four happened simultaneously.
That is a harder story to fit into one traffic-light dashboard. It is also much closer to reality.
The measurement should not punish the truth
That, ultimately, is the principle I take from Spain. Not that enterprise AI governance resembles a pandemic, because it doesn’t. Not that companies are hiding their AI estates, because most aren’t. Not that runtime governance is inadequate, because it solves an important problem extremely well.
The lesson is narrower. A measurement system becomes dangerous when greater visibility can be mistaken for worse performance.
Once that happens, the people who look hardest can appear weakest. The team that finds shadow AI looks less controlled than the team that never searched for it. The company that documents exceptions looks riskier than the company that doesn’t capture them. The organization that expands its inventory can watch its governance percentage fall precisely because its governance programme improved.
That is not a reason to stop measuring. It is a reason to measure the measurement.
Ask what the numerator means. Ask what the denominator contains. Ask who chose it. Ask what sits outside it. Ask whether another source can challenge it. Ask whether finding more truth improves the measurement or worsens the score. And if it worsens the score, ask what behaviour that score will eventually create.
One question for Tuesday morning
The name Spanish flu survived long after historians developed a much richer understanding of the pandemic. That is one of the strange properties of measurements and labels. The first interpretation can become sticky, and by the time better evidence arrives, the story may already have hardened around the earlier number.
AI governance is still young enough that many of its measurements have not hardened yet. We have a chance to decide what we want them to mean.
So this week, I wouldn’t begin by asking your governance team how many AI systems are compliant. I would ask something more basic. How do we know how much AI we have?
Then stay with the answer. Who compiled the inventory? Where did they look? Which sources did they use? What was excluded? What independent evidence could reveal something they missed? How long did they have? And what happens to your governance metrics when they find something new?
Because if discovering another AI system makes your company look worse, while never discovering it leaves the dashboard green, you do not merely have an inventory problem. You have built a measurement system that rewards you for knowing less.
And history has already shown us what can happen when the place that counts most carefully is the place that looks worst.
Idea, thinking and editing by Suneeta Modekurty. Drafted by AI.
Reply to this email with the question your organization uses to decide what belongs in its AI inventory.

