IPM Take
Healthcare AI has become exceptionally good at producing impressive numbers.
Diagnostic accuracy, benchmark scores and comparisons with clinicians can show that a model has capability. They do not show that putting the system into a hospital or virtual clinic improves care.
That distinction is the central argument of a new Nature Medicine Comment from researchers affiliated mainly with Google Research and Google DeepMind, alongside collaborators from Beth Israel Deaconess Medical Center, Stanford and Included Health. The authors argue that trust in conversational medical AI has to be built through prospective evaluation in real clinical environments, because many of the most consequential problems emerge from interactions between the AI, patients, clinicians and healthcare systems rather than from the model in isolation. The Comment was funded by Alphabet, and several authors disclosed Alphabet employment and equity, an important context when interpreting the argument.
The point should matter far beyond Google.
A medical AI that gives the right answer but creates impractical plans, increases clinician workload, works poorly for certain patients or requires unsustainable human supervision has not demonstrated clinical value.
For health systems, payers and HTA bodies, the evidence question is therefore changing from “How well does the model perform?” to “What happens when healthcare actually has to use it?”
Executive Summary
The Nature Medicine Comment argues that retrospective datasets, simulated consultations and benchmark evaluations remain useful for establishing AI capability, but cannot substitute for prospective studies involving real patients and real clinical workflows.
That distinction is already becoming visible in Google’s own research. In a prospective, single-centre feasibility study at Beth Israel Deaconess Medical Center, 100 adults completed a pre-visit text interaction with Google’s research system AMIE before primary-care appointments. A physician monitored each AI interaction and could intervene according to predefined safety criteria. No safety stops were required, and blinded reviewers judged AMIE’s overall differential diagnoses and management plans to be broadly similar in quality to those of primary-care physicians. However, physicians performed better on the practicality and cost-effectiveness of management plans, precisely the kind of difference that conventional diagnostic benchmarking may fail to expose.
The study also had important limitations. It involved one centre, lacked a controlled comparison of the AI workflow against standard care, used live physician supervision, and included only 100 participants. The researchers themselves concluded that it could establish initial feasibility, not prove clinical efficacy or safety at scale.
Google has separately announced plans with Included Health for a nationwide prospective randomized study of conversational AI in real-world virtual care, explicitly moving from simulated performance and small feasibility studies toward controlled evidence at scale.
The broader lesson is not that conversational AI has failed its clinical test. It is that much of that test has barely begun.
Why it matters
- HTA bodies: AI assessment will increasingly need to move beyond technical accuracy toward comparative clinical utility. Relevant outcomes may include patient safety, downstream testing, clinician workload, time saved, access, equity, resource use and whether AI-assisted decisions actually improve care.
- Payers: Paying for an AI system because it performs well on a benchmark risks confusing technical capability with healthcare value. Coverage decisions will need evidence that deployment reduces costs, improves access or outcomes, or releases clinical capacity without creating expensive downstream consequences.
- Industry / innovation partners: Benchmark leadership may help demonstrate that a model is technically competitive, but it will become harder to use benchmark performance as the primary proof of clinical readiness. Companies that can generate prospective evidence inside real workflows may gain a more meaningful advantage than those competing only on model scores.
Healthcare has seen this story before.
A technology performs impressively under controlled conditions, generates enthusiasm and reaches clinical practice, only for the real system to reveal complications that were invisible in development.
Conversational AI makes that problem particularly difficult because the “intervention” is not static. The system interacts with different patients, receives incomplete information, responds to unpredictable questions and sits inside workflows involving clinicians, electronic records, appointments, diagnostic testing and reimbursement.
That complexity is why the new Nature Medicine Comment draws a hard line between demonstrating capability and demonstrating clinical value.
Laboratory studies remain important. Researchers have shown that large language models can perform strongly on clinical reasoning tasks, simulated diagnostic encounters and increasingly sophisticated multimodal consultations. Those experiments help answer whether a model can perform particular tasks. They are far less capable of showing what happens once the AI becomes one component of a healthcare system.
The clinic exposes what the benchmark cannot
Google’s prospective AMIE study illustrates the difference unusually well.
The system interacted with 100 patients before ambulatory primary-care appointments, taking histories and producing summaries for clinicians. Human supervisors watched the conversations live and were prepared to stop an interaction for potential clinical harm, significant emotional distress, risk of harm to self or others, or a patient’s request to end the session. None required intervention.
That sounds encouraging, but the more revealing finding came later.
When clinical evaluators compared AMIE with the treating physicians, overall differential diagnosis and management quality was similar. Yet physicians produced management plans judged more practical and cost-effective. AMIE did not have access to the electronic health record, could not perform a physical examination and lacked some of the local clinical context available to physicians.
Those limitations are not peripheral.
They are the healthcare system.
An AI can recommend a clinically defensible investigation while still recommending something unavailable locally, unnecessarily expensive or poorly aligned with the patient’s actual pathway. A benchmark that rewards the medically plausible answer may never capture that failure.
The same applies to safety. Zero safety stops across 100 supervised interactions is useful feasibility evidence. It is not evidence that an unsupervised system will remain safe across millions of conversations, languages, disease severities and patient populations.
Prospective research forces those distinctions into the open.
The evidence standard may become the real competitive advantage
This has major implications for the AI market.
Healthcare organisations are increasingly being offered systems described as clinician-level, expert-level or capable of improving access. Those claims can be grounded in legitimate research while still leaving unanswered questions about deployment.
How much supervision is required? Which patients should not use the system? Does it create extra investigations? Does it reduce clinician time after the time spent reviewing its output is counted? Does it work for people with limited health literacy? What happens when the model changes six months after the study?
These are not ordinary software-performance questions. They are implementation and clinical-effectiveness questions.
Google’s February announcement of a planned nationwide randomized study with Included Health is therefore significant less because of any expected result and more because of the study design itself. The stated goal is to compare conversational AI with standard virtual-care workflows using real patients at scale rather than relying only on simulations.
If that model of evaluation spreads, the competitive landscape for medical AI could change.
The strongest product may no longer be the model with the highest benchmark score.
It may be the one with the strongest evidence that patients, clinicians and health systems are actually better off when it is switched on.
That evidentiary shift is particularly important for precision medicine, where AI is increasingly expected to interpret complex patient information, personalise recommendations and expand access to expertise.
Precision without prospective validation risks becoming personalised uncertainty at scale.
Clinical AI does not need less innovation.
It needs evidence capable of keeping up with it.

