General-purpose LLMs beat FDA-cleared clinical AI in a head-to-head, exposing a validation gap
Original reporting: Nature Medicine
A June 2026 Nature Medicine benchmark found frontier general models outperformed FDA-cleared clinical AI tools on real physician queries, raising a question regulators have not answered: does a cleared device actually beat the free model a clinician already has?
Why it matters
The benchmark did something most vendor studies avoid: it compared cleared clinical tools against the general models a clinician could open for free, on real questions.
The result reframes the whole conversation. The useful bar is not whether a tool clears a threshold in isolation, but whether it beats the alternative already sitting on every desk.
The ReasonFirst take
Clearance tells you a tool cleared a bar. It does not tell you it beats what you already have in another tab. Ask for the head-to-head, not the badge.
Who should care
What to watch
Whether procurement and regulators begin requiring comparison against a general-model baseline, not just a fixed accuracy threshold.
A question worth sitting with
When the free tool matches the cleared one, what exactly is the clearance certifying?
More signals
MDCalc adds quality ratings to its 800-plus clinical calculators
STAT News · July 17, 2026
OTC Continuous Glucose Monitors Are Now Available for Toddlers. The Evidence for Obesity Use Is Not.
STAT News · July 8, 2026
Diagnostic test overuse is a structural problem, not a willpower problem
STAT News · July 6, 2026