Medical AI now matches or beats doctors in controlled tests, but is not ready to replace clinical decisions
Two medical AI systems, Mira and AMIE, showed strong results in diagnosis and clinical management. The progress is meaningful, but the studies were conducted in controlled settings and still leave open questions about real-world hospital use.
Summary
Two medical artificial intelligence systems, called Mira and AMIE, have shown strong results in controlled tests of diagnosis and clinical management.
According to the Financial Times, Mira, developed by researchers in Germany, reached 87.1% diagnostic accuracy across eight emergency conditions, above the 78.1% achieved by a panel of six specialist physicians. The system uses electronic health record data and can choose from thousands of clinical options, including tests, medication, and procedures.
AMIE, developed by Google and based on Gemini, was tested in multi-visit simulated primary care scenarios. In the described study, it produced treatment and investigation plans that were more precise and more aligned with UK clinical guidelines than those produced by primary care physicians.
In practice
The most important reading is not that AI has beaten doctors. That conclusion would be too simplistic.
What these studies suggest is that specialized medical AI systems may support diagnosis, triage, disease management, and treatment planning in well-defined contexts. They can act as a second layer of clinical reasoning: organizing information, suggesting hypotheses, checking guidelines, and reducing omissions.
For AMIE, the tests involved actors role-playing patients and structured case scenarios. For Mira, clinical cases were also passed through simulations with AI agents acting as patients. That makes the results interesting, but still different from real clinical practice, where there is noise, incomplete data, time pressure, difficult communication, and legal responsibility.
What we still don't know
The biggest risk is confusing performance in simulation with clinical readiness. A tool can perform well in controlled tests and still fail when deployed in a hospital, a real consultation, or a health system with inconsistent data.
There is also a strategic question: specialized medical models may age quickly if general-purpose models improve rapidly. If systems such as Gemini, Claude, or GPT become strong enough in clinical reasoning, guideline use, and tool operation, part of the advantage of purpose-built medical models may shrink.
That does not make specialization useless. In medicine, validation, safety, traceability, clinical data integration, and accountability remain essential. But it suggests the future may not be just a closed medical model, but a combination of strong general models, validated clinical layers, and physician supervision.
Why it matters
- Medical AI is moving from generic demonstrations to more concrete clinical tasks.
- The results show potential to support doctors, but do not justify replacing human supervision.
- Real-world validation remains the difference between promising research and safe clinical use.
- The future value may come less from AI that gives answers and more from auditable systems that help doctors make better decisions.