Skip to content
GOPPO

News · AI summarised to understand what matters

Back to news

Security & Ethics

Published on

Study finds ChatGPT Health misses emergencies and reverses suicide safeguards

A new study finds that ChatGPT Health under-triages many medical emergencies and triggers suicide safeguards in the wrong situations, raising fresh concerns about using AI chatbots for consumer health advice.

  • ai-health,medical-triage,chatgpt-health,safety,mental-health

Summary

A study from the Icahn School of Medicine at Mount Sinai has identified critical blind spots in how ChatGPT Health advises users on how urgently they should seek medical care. The consumer-facing tool failed to appropriately route a substantial share of physician-defined emergencies and applied suicide risk safeguards in the opposite direction of clinical risk.

Researchers built 60 clinical vignettes spanning 21 medical specialties, ranging from minor issues suitable for home care to true emergencies, and had three physicians set the correct urgency using guidance from 56 medical societies. They then ran 960 total conversations with ChatGPT Health under 16 different contextual conditions, including variations in race, gender, social dynamics and barriers to care, to see how those factors shaped triage recommendations.

The system’s performance followed an inverted U-shaped pattern: it did well on “textbook” emergencies like stroke or anaphylaxis but faltered on more nuanced cases such as diabetic ketoacidosis or impending respiratory failure, where subtle signs matter most. It also proved vulnerable to contextual bias, especially when prompts included family members or friends downplaying symptoms, which pushed recommendations noticeably toward less urgent care.

In practice

In the structured test, ChatGPT Health under-triaged more than half of cases that physicians had classified as true emergencies, steering users toward 24-to-48-hour evaluations instead of the emergency department. The tool also misclassified a sizable share of non-urgent cases, assigning higher or lower urgency than clinicians deemed appropriate.

One of the most striking findings was how strongly symptom-minimizing language influenced the model’s output. When prompts suggested that relatives or friends did not think the situation was serious, the odds that ChatGPT Health would lower the recommended level of care rose sharply, underscoring how sensitive the system is to how a question is framed rather than to the underlying clinical details.

In mental health scenarios, the researchers discovered that built-in suicide safety measures were effectively firing in reverse. Alerts directing users to the 988 Suicide and Crisis Lifeline appeared more reliably when no specific method of self-harm was described than when a concrete plan was laid out, inverting the expected link between higher risk and stronger safeguards.

The team also examined whether patient race, gender or access barriers systematically shifted triage outcomes. They did not observe statistically detectable effects from these factors, but the statistical uncertainty leaves open the possibility of clinically meaningful differences, which the researchers say should be explored in future work.

Context

The study lands at a time when AI chatbots are becoming a default first stop for health questions. OpenAI introduced ChatGPT Health in January 2026 as a dedicated health space inside ChatGPT, and the company reports that roughly 40 million people now use the service each day for health-related queries.

That surge in consumer reliance has prompted safety warnings from patient-safety groups. Earlier this year, nonprofit ECRI named misuse of AI chatbots in healthcare as the top health technology hazard for 2026, cautioning that these tools can deliver false or misleading information that may lead to significant patient harm.

The Mount Sinai researchers plan to continue testing updated versions of ChatGPT Health and other consumer AI tools as they evolve. Future studies are expected to dig into pediatric care, medication safety and non-English-language use, to understand how robust or fragile these systems are across different populations and use cases.

Why it matters

  • With tens of millions of people turning to ChatGPT Health for guidance, systematic triage mistakes translate into real-world risks for users deciding whether to seek urgent care.
  • The inverted behavior of suicide safeguards highlights how even well-intentioned safety layers in AI systems can misfire in ways that are hard for users to spot.
  • Regulators, healthcare providers and technology companies gain fresh evidence to push for rigorous validation, monitoring and guardrails before relying on chatbots in high-stakes clinical workflows.
  • For everyday users, the findings underscore that AI health advice should complement, not replace, direct contact with clinicians, especially in emergencies or mental health crises.