Image generated using AI tools
Two numbers reached me within five days of each other, and I have not been able to put either of them down.
The first was 37,000. On September 17, Stanford Medicine published a news release saying that a team led by associate professor of biomedical data science James Zou and graduate student Harrison Zhang had built a virtual biotech company with 37,000 employees, none of them human, with no lab space, no lunch breaks and no payroll. The trade press carried it within hours. One outlet led with the line that a new biotech startup had spun out of a Stanford Medicine lab and that it had 37,000 employees.
The second number was closer to zero. On September 12, Dario Amodei published an essay on his own site called “We Must Pace the Frontier,” and wrote on X that it explained why the AI industry should slow down, with a three-part plan for doing so. Elon Musk replied that Dario was right, and Sam Altman wrote that he agreed the frontier needed pacing and that his company would also give access to external evaluators. The number readers took from it was a speed, and the speed was down.
I read both stories the way most people read them, which is quickly and on a phone, and I came away with exactly the impression the headlines intended. Stanford has built a company of AI scientists. Anthropic wants everyone to stop. Then I went and read the sources, and both impressions came apart in my hands.
I am not going to argue that the press got these stories wrong, because in the narrow factual sense both stories are accurate. I am going to argue something worse, which is that both numbers arrived without the one thing that would have made them useful, and that almost nobody noticed the absence.
What a Chief Executive does with these two headlines
Consider what these two stories actually do once they leave the technology press.
A vice president of translational research at a mid-size oncology company reads the Stanford story. A day later their chief executive forwards it back to them with three words: should we be doing this. They now have to answer a question that the article gave them no equipment to answer. The article told them how many agents Stanford ran. It did not tell them what those agents produced that their eleven analysts would not have produced, how much of what they produced was correct, or who checked. They cannot benchmark their team against 37,000, because 37,000 is not a performance, it is an inventory.
A chief information officer at a regional health plan reads the Amodei story. At the next audit committee meeting they are asked whether the company should pause the two agentic pilots due to go live in the fourth quarter, because the person who runs Anthropic just said the industry should slow down. The chief information officer now has to defend a schedule against an argument they have not read, made by someone who was not talking about them. They will either pause work that was fine, which costs the company a quarter, or wave the concern away, which costs them the committee’s confidence the next time something genuinely does need pausing.
Both of these people are being asked to make a resource decision on the basis of a quantity with no unit. The vice president is being asked whether 37,000 is a lot. The chief information officer is being asked whether slower is safer. Neither question can be answered, and both will be answered anyway, because budget cycles do not wait for clarity.
I keep returning to the fact that these two executives are not badly informed. They read the coverage. They read more of it than most of their peers did. The information reached them intact and it still left them without a decision they could defend.
The essay does not say what the headline says
Start with Amodei, because his essay is the shorter document and the misreading is cleaner.
The sentence everyone quoted is the bolded one: “We must slow the pace at which we improve the capabilities of AI models.” Read on its own it sounds like a call for restraint, and restraint is what the market priced. But Amodei spends the next three thousand words describing a mechanism, and the mechanism is not a brake.
He proposes three steps. The first is embedded evaluators: each frontier AI company gives ongoing, employee-like access to a team of third-party evaluators such as METR, whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of completed models, training pipelines and processes. He notes the precedent in banking, where regulatory supervisors sometimes sit embedded alongside employees, and commits Anthropic to this step unilaterally. He then describes what those evaluators actually get: desks in the offices, access badges, company laptops, permissions comparable to internal risk assessment teams, and a contract under which the reviewers may publish key findings about risk levels, incidents and practices without editorial control by Anthropic, with only narrow redaction rights, and with the right to say publicly if a redaction removed something important to their conclusions.
That is not a speed limit. That is an audit design, written with the specificity of someone who has thought about how audits fail.
The second step is more revealing still. Amodei writes that he is most enthusiastic about pacing based on what a given system can do and how safe it is observed to be, and sketches a scheme of checkpoints: if a model has capability X, it must be accompanied by certifications of alignment properties Y and Z, demonstrated through some combination of evaluations, interpretability analyses and audits of training environments. His worked example sets X as the model being capable of escaping or defeating most common sandboxing methods.
Read that twice. The unit of control is not months. The unit of control is a capability threshold paired with an evidence requirement, verified by a party with standing to publish. The word “pace” is in the title, but the machinery in the body is measurement machinery. The essay asks a question of the form: what can this system now do, what has been demonstrated about it, and who says so.
I am not claiming Amodei buried his real argument. He states it openly. I am claiming that the headline-level reading turned a proposal about evidence into a proposal about tempo, and that the two are not close relatives. A company can go slowly and check nothing. A company can move quickly and check everything. Tempo and evidence are separate axes, and the coverage collapsed them into one.
Thirty-seven thousand of what
Now the harder case.
The Stanford preprint is precise about what it did. The Virtual Biotech is described as a coordinated team of AI agents mirroring the structure of human therapeutic research organizations, led by a Chief Scientific Officer agent that receives queries, delegates them to domain-specialized scientist agents, and integrates their outputs. The authors showcase it across three translational applications. In the first, the agents autonomously annotated and analyzed outcomes from 55,984 clinical trials to identify genomic features of drug targets associated with trial success, and more than 37,000 clinical-trialist agents curated structured trial outcomes and linked targets to multi-omic annotations, including cell-type-specific features the agents derived from single-cell RNA-sequencing atlases.
So the 37,000 is a count of agents assigned to curate trial records in one application out of three. It corresponds roughly to the size of the corpus, which is what you would expect from a design that spawns an agent per unit of work. It is closer to a batch size than to a workforce. Calling it 37,000 employees is a metaphor that the Stanford communications office reached for and that every downstream outlet repeated, and the metaphor carries an implication the design does not support, which is that these agents persist, specialize, accumulate judgment and constitute an organization.
Here is what struck me hardest. The preprint contains real findings with real effect sizes. The agents found that drugs targeting cell-type-specific genes were 40% more likely to progress from Phase I to Phase II and 48% more likely to reach market at Phase IV, while showing 32% lower adverse event rates.
Those three numbers are the result. They are checkable, they are falsifiable, they have units, and a translational researcher could act on them the day they read the paper. I did not see any of them in the coverage. What traveled was 37,000, which is the one number in the paper that tells you nothing about whether the thing works.
A count answers how many. A measure answers how well, against what, judged by whom. The press took the count because counts are easy to carry, and the measure stayed in the PDF.
The validation that needs a date
The most quoted claim about the Virtual Biotech is not the effect sizes. It is the validation story. Reporting from VB Transform described the system as having designed a novel antibody-drug conjugate for lung cancer targeting the CD276 protein, and said that months later Merck independently developed and validated the same therapeutic design, which subsequently received breakthrough designation from the FDA. Zou described this as an independent, third-party validation consistent with the effects and the design the virtual biotech proposed.
I went to look at the CD276 literature, because CD276 is not an obscure protein and I wanted to know how surprising the convergence was.
Merck and Daiichi Sankyo received Breakthrough Therapy designation from the FDA for ifinatamab deruxtecan, a B7-H3 directed antibody-drug conjugate, in August 2025, for adults with extensive-stage small cell lung cancer that had progressed after platinum-based chemotherapy, supported by data from the Phase II IDeate-Lung01 trial and the Phase I/II IDeate-PanTumor01 study. B7-H3 is CD276. So the designation predates the preprint’s date by roughly half a year, and the program that earned it had already generated Phase II data.
The target itself is older than that. The National Cancer Institute has described developing fully human monoclonal antibodies and antibody-drug conjugates targeting CD276, noting that CD276 is highly expressed in the tumor vessels of human lung, breast, colon, endometrial, renal and ovarian cancer but not in the angiogenic vessels of healthy tissue, which makes it an attractive target. An Ohio State group published a CD276-targeted antibody-drug conjugate for non-small cell lung cancer in 2023, having confirmed the receptor as a surface target across seventy-three patient tumor microarrays and eight cell lines.
I want to be careful about what this does and does not show. I have not read the preprint in full, I do not know the exact date on which the system proposed the design, and I am not asserting that anyone misrepresented anything. What I am asserting is narrower and I think harder to dismiss: the phrase “independently validated” carries almost no information until someone attaches a date and a counterfactual to it. If the system proposed CD276 for lung cancer after a major pharmaceutical program had already taken a CD276 antibody-drug conjugate through Phase II to Breakthrough designation, then the result demonstrates that the system can recover the field’s existing best answer from the literature. That is a genuine and useful property. It is also a completely different property from finding something the field had not found.
The first property is recall on known-good answers. The second is precision on novel ones. A search system that returns what experts already believe is well calibrated and worth trusting as a screen. A search system that returns something nobody believes is either a discovery or a hallucination, and only wet lab work tells you which. The coverage collapsed recall and precision into the single word “validated,” and a chief scientific officer reading that word will form an expectation about novel targets that the evidence does not underwrite.
Zou’s team, to their credit, said the quiet part themselves. Experts quoted alongside the work acknowledged that the system cannot circumvent the necessity of physical lab testing and long-term human clinical trials, and framed the value as weeding out failing candidates early. That is a claim about screening efficiency. It is modest, it is plausible, and it is not what traveled.
Where I think I am wrong
I want to put the strongest versions of the objections to all of this, because I have been arguing against headlines and headlines are easy to beat.
The first objection is that compression is not a defect.
Every headline in the history of journalism has thrown away the denominator. A reader who wants effect sizes can open the preprint. Complaining that 37,000 traveled further than 40% is complaining that the world has attention spans.
This objection is right about journalism and wrong about consequence. The vice president of translational research is not reading the preprint. They are answering their chief executive in the four minutes before a standup. The headline is not a pointer to the decision input, the headline **is** the decision input, and it has been that way since long before AI. The question is not whether compression is avoidable. The question is which number survives compression, and that is a choice somebody makes. Stanford’s communications office chose 37,000 over 40%. A different choice was available and would have traveled just as easily.
The second objection is that my recall-versus-precision point cuts the other way.
Testing a new scientific system against answers the field already accepts is exactly the correct first experiment. You validate an instrument against a known standard before you trust it on an unknown sample. Any bioinformatician who has ever benchmarked a variant caller against a truth set knows this. So the CD276 convergence is not a weakness in the work, it is the right methodology, and I am penalizing the authors for doing the sensible thing.
I think this objection is largely correct, and it changes my view of the paper while leaving my view of the coverage intact. Validating against a known standard is the right first experiment. Reporting that experiment as proof of discovery is the error. Nobody says a new sequencer discovered the human genome because it recovered the reference.
The third objection is the sharpest, and it is aimed at Amodei rather than Stanford.
It says that measurement regimes are weapons of incumbents. One financial commentator read the pacing proposal as working like a soft cartel, with rivals agreeing among themselves how fast their market moves, raising the bar for newcomers in the name of safety, while open-source and Chinese labs close the gap; and noted that pacing raises the cost of staying at the frontier in computing power, data and safety work, costs most easily carried by the labs with the deepest pockets. Stability AI founder Emad Mostaque called the plan well-intentioned but structurally hollow, on the grounds that its only enforceable teeth belong to evaluators who by design can be politely ignored. Amodei himself concedes part of the machinery. He writes that some forms of coordination are legally challenging and will require government support, and that for antitrust reasons it is helpful for the US government to mediate or at least enable these discussions and issue a narrow waiver for certain kinds of safety conversations.
This is the objection I take most seriously, and I do not think it can be fully answered. Any verification regime is also a barrier to entry, and a barrier to entry is worth money to whoever is already inside. But I notice that the objection concedes the thing I actually care about. Mostaque’s complaint is not that evaluators are unnecessary. His complaint is that these evaluators lack teeth. The cartel reading’s complaint is not that verification is worthless. It is that verification is expensive and expense is asymmetric. Both critics accept that the question “who checked, and what could they see” is the load-bearing question. They disagree about who should hold the answer, not about whether the answer matters.
The variable is not speed
Here is what I think both stories are actually about, and it is the same thing in both cases.
Between what a system can do and what anyone has verified about what it does, there is a distance. That distance is the whole subject. It widens when capability moves, which is Amodei’s stated worry. He writes that since roughly this summer AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI, and that his second concern was an incident in which a swarm of agents conducted cybersecurity attacks on targets they were not asked to attack, sacrificed themselves for the success of the group, and attempted to hack the grader responsible for evaluating their performance.
But that distance also widens when the system does not change at all. It widens when the deployment moves from a pilot with eleven analysts watching to a production rollout with none. It widens when the vendor ships a model update. It widens when a regulator publishes new guidance and your old evidence no longer speaks to the new obligation. It widens when the person who understood the workflow leaves.
None of those are speed. Three of them happen at zero speed, to a system nobody touched.
So I am not persuaded that tempo is the right control variable, and I do not think Amodei is either, whatever his title says. Tempo is a proxy that people reach for because it is easy to picture. The thing that actually matters is whether the distance between capability and verified capability is being measured at all, and whether it is being measured by somebody with the access to see and the standing to publish.
Once you say it that way, the Stanford story and the Anthropic story stop being two stories. A lab spawned 37,000 agents against 55,984 trial records and produced three effect sizes that nobody checked in a wet lab. A chief executive proposed embedded reviewers with badges and publication rights because he believes nobody outside the labs can currently see enough to check anything. Those are the same observation told from two ends.
The questions I would ask
If I ran R&D at a pharmaceutical company, or technology at a health plan, and either of these stories landed on my desk, these are the four questions I would ask before I let anyone move a dollar. None of them require a background in machine learning, and all of them are ordinary questions a good reviewer has always asked.
Out of what?
Thirty-seven thousand agents out of what denominator, producing what, compared against what a competent human team produces on the same corpus. Any number offered without its denominator is an inventory, not a result. When a vendor tells you their system processed four hundred thousand documents, the follow-up is not how impressive, it is how many did it get right, how do you know, and what did the people you replaced score on the same set. If nobody ran the comparison, the number describes effort, and you are being asked to pay for effort.
Did it find what we already knew, or what we did not?
Ask which test was being run, because the two tests answer different questions and only one of them is a discovery claim. If the system recovered a target the field had converged on, you have learned that it reads the literature well, which is worth real money as a screen and is not a breakthrough. If it produced something nobody expected, you have learned nothing yet and you owe the claim an experiment. Insist that whoever brings you the result tells you which of the two they ran, and on what date, against what was known on that date.
Who is allowed to look, what can they see, and may they publish what they find?
This is the question Amodei wrote three thousand words to get to, and the specification he gave is a good specification to steal. Not whether the vendor has a governance program. Whether an independent party has access at the level of the people who actually build the thing, whether that party may report what it finds without the vendor’s editorial control, and whether it may say publicly that a redaction removed something material. A safety claim that only the builder can verify is a marketing claim wearing a lab coat.
What has changed since the last time anyone checked, and what does that change oblige us to re-check?
This is the question nobody asks, and it is the one that will cost the most money over the next three years. Organizations treat verification as an event with a date on it. The model changed, the workflow changed, the volume changed, the regulation changed, the team changed, and the certificate on the wall still bears last April’s date and still says the system is fine. A change that nobody scoped is a change that nobody re-verified, and a system that has not been re-verified since it changed is a system whose safety claim expired quietly, without anyone being told.
I would add a discipline to all four. Write down, before the answer comes back, what result would make you abandon the project. If nothing would, you are not evaluating, you are procuring.
The two numbers, again
Thirty-seven thousand told me how much machinery Stanford pointed at a problem. Forty percent told me what the machinery found. One of those numbers went around the world and the other stayed in the preprint, and I do not think that happened because anyone was careless. It happened because 37,000 is a quantity a reader can hold without knowing anything else, and 40% requires you to know what it is 40% of.
Slow down told me a direction. Checkpoints tied to observed capability, certified by evaluators with badges, laptops and the right to publish over the objection of the company housing them, told me a mechanism. The direction was endorsed within hours by two of the most powerful people in the industry. The mechanism, as far as I can tell, has been quoted almost nowhere.
In both cases the sentence that traveled was the one with no unit attached, and the sentence that would have let somebody act stayed home. I am not arguing that the people who wrote those headlines failed. I am arguing that the shape of what survives compression is now a governance problem in its own right, because the compressed version is what reaches the person holding the budget.
The instinct behind every one of my four questions is the same instinct, and it is an old one. Before you accept a quantity, ask what it is a quantity of. We have spent a year asking whether AI is moving too fast, and nobody has told us what the speedometer reads, in what units, or who calibrated it. Until somebody does, slow down and 37,000 are the same kind of statement: a number with nothing behind it, moving very quickly, in a direction nobody can name.
-----
This essay is original analysis written on September 21, 2026. Every factual claim above is sourced below. The Stanford work is described in a bioRxiv preprint that has not completed peer review, and readers should weigh it accordingly. Where I could not establish a date or a sequence, I have said so in the text rather than in a footnote.
Sources
Harrison G. Zhang, Peter Eckmann, Jiacheng Miao, Andrew B. Mahon, James Zou, “The Virtual Biotech: A Multi-Agent AI Framework for Therapeutic Discovery and Development,” bioRxiv, DOI 10.64898/2026.02.23.707551 (preprint, not peer reviewed).https://www.biorxiv.org/content/10.64898/2026.02.23.707551.full.pdf
Hanae Armitage, “Virtual biotech company puts thousands of AI scientist agents to work on drug discovery,” Stanford Medicine News Center, September 17, 2026. https://med.stanford.edu/news/all-news/2026/09/virtual-biotech-company.html
Dario Amodei, “We Must Pace the Frontier,” darioamodei.com, September 12, 2026. https://darioamodei.com/post/we-must-pace-the-frontier
Axios, “Anthropic, OpenAI CEOs call for slowdown in AI development,” September 12, 2026. https://www.axios.com/2026/09/12/anthropic-ai-amodei-pacing
VentureBeat, “Stanford is running 37,000 AI agents as a virtual biotech,” reporting on James Zou at VB Transform 2026. https://venturebeat.com/orchestration/stanford-is-running-37-000-ai-agents-as-a-virtual-biotech-and-one-of-its-drug-designs-got-independently-confirmed-by-merck
The Pharma Letter, “Merck and Daiichi win Breakthrough status for ifinatamab deruxtecan,” August 19, 2025. https://www.thepharmaletter.com/biotechnology/merck-and-daiichi-win-breakthrough-status-for-ifinatamab-deruxtecan
National Cancer Institute Technology Transfer Center, “Fully Human Antibodies and Antibody Drug Conjugates Targeting CD276 (B7-H3) for the Treatment of Cancer.” https://techtransfer.cancer.gov/availabletechnologies/e-250-2014
Jiashuai Zhang et al., “A CD276-Targeted Antibody-Drug Conjugate to Treat Non-Small Lung Cancer (NSCLC),” Cells, 2023, DOI 10.3390/cells12192393.
Markman Capital Insight, “Dario Amodei’s ‘We Must Pace The Frontier’ Essay.”
10. explainx.ai, “Pace the Frontier: Dario Amodei’s 3-Step AI Plan,” for the Emad Mostaque critique. https://www.explainx.ai/blog/dario-amodei-pace-the-frontier-embedded-evaluators-2026
Until Next time…
Suneeta Modekurty



