The Model Isn’t the Medicine
Nature Medicine recently published a rigorous independent evaluation of clinical AI. Researchers at NYU compared three frontier large language models (GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6) against two purpose-built clinical AI tools, OpenEvidence and UpToDate Expert AI, across three evaluation stages: 500 standardized medical knowledge questions, 500 clinician-alignment items, and 100 real clinical queries (RCQs) drawn from physician use at NYU Langone Health, reviewed by twelve blinded clinicians.
The finding was decisive, and the response from the medical community was explosive. That’s because frontier models outperformed purpose-built clinical tools across every dimension. On real-world clinical queries, frontier LLMs scored roughly 10 to 14 percent higher in aggregate. The specialized tools performed comparably to auto-enabled Google Search.
I am an engineer who has spent my career building smart systems, not a doctor. From where I sit, independent, rigorous evaluation of clinical AI has been rare enough that its arrival warrants genuine attention. Health systems, procurement offices, and regulators should read this paper carefully. The evidence that frontier general-purpose models outperform purpose-built clinical tools on real physician queries is consequential.
About the author
Matthew Rothstein
Matthew Rothstein is the Head of Engineering at Akido, an AI-native healthcare company focused on expanding access to care through technology. He leads the engineering organization responsible for building and scaling the technology that powers Akido