General AI models outperform clinical tools in benchmark
A study compared specialized clinical tools with GPT-5, Gemini 3 Pro, and Claude Sonnet 4.5. Its finding raises a practical question: a medical label on an AI model does not guarantee better performance.
Summary
A study published on arXiv compared specialized clinical tools, including OpenEvidence and UpToDate Expert AI, with recent general-purpose models such as GPT-5, Gemini 3 Pro, and Claude Sonnet 4.5.
The evaluation used a 1,000-item mini-benchmark combining medical knowledge tasks, including MedQA, with clinician-alignment tasks such as HealthBench. According to the authors, the general-purpose models consistently performed better, with GPT-5 achieving the highest scores.
In practice
The conclusion is not that doctors should be replaced by general-purpose chatbots. It is that clinical tools marketed as specialized need more transparent evaluation.
In medicine, “specialized” naturally sounds safer. But if frontier general-purpose models improve faster, with more data, better reasoning, and stronger safety mechanisms, they may outperform more closed or less frequently updated clinical tools in some scenarios.
For hospitals, healthcare professionals, and companies building clinical software, this changes the question. Instead of asking only “was this model trained for medicine?”, the better question is “has it been tested against the best available models, on relevant tasks, with public criteria?”.
What we still don't know
Benchmarks are not real clinical practice. Answering medical questions or simulated cases well is not the same as dealing with patients, incomplete exams, legal responsibility, and hospital decisions.
It also remains unclear how these results translate into clinical environments with private data, real workflows, and medical supervision. The prudent conclusion is that recent general-purpose models should be included in comparisons, not that they should replace clinical tools or healthcare professionals.
Why it matters
- It challenges the idea that AI models labelled as medical or clinical are always superior.
- It pressures clinical AI tools to publish independent and comparable evaluations.
- It shows that the pace of general-purpose LLM development may overtake closed vertical products.
- It reinforces the need for real clinical validation before using AI in medical decisions.