Stanford researchers say they have built an AI-powered “virtual biotech” organization of 37,000 specialized software agents to take on parts of drug discovery that are ordinarily divided among scientific teams: choosing targets, examining safety signals, reviewing clinical-trial histories and sketching future studies. Stanford Medicine said Sept. 17 that the work, led by biomedical data scientist James Zou and graduate student Harrison Zhang, was published in Science.
The scale is notable, as are the team’s reported findings from about 50,000 clinical trials. But the evidence described so far is computational and retrospective. The system has not produced a medicine tested in people, and Stanford says the next stage is to examine additional findings in physical laboratories. The difference is substantial in an industry where a plausible target, a drug candidate and a clinically successful treatment are separated by years of experiments, toxicology work and trials.
A software organization rather than a biotech company
Stanford’s description uses the language of a company, but this is not a conventional startup with employees, laboratories or a drug pipeline. It is a coordinated multi-agent AI system organized into divisions, overseen by a chief science officer agent. The agents are assigned specialized work and exchange outputs as the system moves across tasks in a drug-development workflow.
That design extends Zou’s earlier virtual-lab work, which enlisted collaborating AI scientist agents for in-silico research. In Stanford’s June account of its Health AI week, Zou had already described a virtual-biotech effort using 37,000 agents to read and organize clinical-trial material. A 2025 Columbia University seminar listing also documented related Virtual Lab work, including prior nanobody research that Stanford had said received experimental validation. That earlier context is useful, but it is not independent confirmation of the new virtual-biotech results.
Multi-agent systems are increasingly promoted as a way to make general-purpose language models more useful for complicated knowledge work. Rather than asking one model to reason through every step, developers split a problem among agents with distinct roles, then combine or challenge their conclusions. The arrangement can improve organization and breadth, but it can also amplify errors when many agents draw on incomplete, inconsistent or biased records. A large agent count is therefore a measure of workflow design and computing ambition, not by itself a measure of scientific reliability.
What the trial analysis found
According to Stanford, the agents catalogued and analyzed roughly 50,000 clinical trials in less than a week. The group used trial and molecular information to examine two properties of prospective drug targets: how specifically a target is concentrated in a desired cell type, and whether its gene expression is bimodal—more like an on-off switch across cells than a broadly distributed signal.
Stanford reports that drug programs scoring highly on both measures were 40% more likely to move from phase 1 to phase 2, 48% more likely to reach market, and associated with 32% fewer adverse events than drugs with broad activity. The institution said the patterns appeared across cancer as well as brain, heart, kidney and lung conditions.
Those numbers are associations from a retrospective analysis, not evidence that selecting a target with those characteristics will cause a drug to succeed. The available account does not provide the number of programs in each comparison, patient-level data, uncertainty ranges, adjustment methods or details on which trials had usable single-cell molecular information. Those details are necessary to assess how robust the reported relationships are and whether they generalize to new targets.
Zou’s proposed biological explanation is plausible as a hypothesis: a target concentrated in a particular cell type and expressed in a switch-like pattern may offer a clearer therapeutic window than one active across many cells. Yet drug safety and efficacy also depend on the treatment modality, dose, delivery, disease biology and effects that cannot be inferred cleanly from gene-expression patterns. The system may be useful for prioritizing questions, but it does not erase those experimental constraints.
A retrospective B7-H3 comparison, not clinical proof
To demonstrate the system’s design capability, the team gave it information available before January 2025 and asked it to develop a strategy for lung cancer. Stanford says the agents proposed an antibody-drug conjugate aimed at B7-H3, a cell-surface protein used as a target in cancer-drug research. Antibody-drug conjugates combine an antibody that seeks a selected target with a cell-killing payload.
Stanford says an unnamed pharmaceutical company independently developed a similar B7-H3 antibody-drug conjugate strategy in August 2025 and that the therapy later received FDA breakthrough therapy designation. The timing is intended as a form of retrospective comparison: the system’s proposal was generated using a knowledge cutoff that preceded the company’s reported program.
It is not, however, a head-to-head prospective validation. Stanford did not identify the company or therapy, so the match in target, payload, patient population, molecular design and clinical evidence cannot be checked from the public account. Nor does breakthrough therapy designation mean FDA marketing approval; it is an expedited-development designation for a program addressing a serious condition where preliminary evidence suggests it may offer a substantial improvement over available therapy.
The project’s more immediate value may lie in triage. Drug developers face far more potential targets and trial configurations than they can test at the bench or in the clinic. A system able to trace evidence across thousands of trials could help scientists find patterns worth challenging with experiments, flag possible safety concerns or identify trial designs that deserve closer inspection.
That is a narrower claim than an autonomous drug-discovery company, but it is also the one supported by the work described so far. Stanford’s stated plan to test new findings in laboratories will determine whether the 37,000-agent organization can do more than efficiently generate hypotheses from the record of drug development.
