Particle.news

Warentest Finds Mixed Medical Performance From AI Chatbots, ChatGPT Leads

The consumer tests show some chatbots can give useful first assessments while gaps in data handling and variable accuracy mean they cannot replace medical care.

Overview

  • Stiftung Warentest tested five free, general-purpose chatbots — ChatGPT, Gemini, Meta AI, Claude and DeepSeek — using five clinical model cases and found three produced acceptable first assessments.
  • ChatGPT produced correct suspected diagnoses and appropriate action advice in all five scenarios in the review, while Gemini performed worst and once ended a conversation with “I cannot help you.”
  • The testers highlighted serious privacy concerns because providers may store or process health data on servers outside Europe and advise using chatbots without logging in or switching to dedicated symptom‑checkers for regular use.
  • All five bots now give responsible responses to suicide prompts by urging professional help and refusing to provide methods, and experts warned diagnostic accuracy still depends heavily on complete user input such as age, sex and medical history.
  • Independent studies showing error rates and a U.S. lawsuit over a missed emergency diagnosis are increasing scientific and legal scrutiny that could push for clearer safety rules, transparency and vendor liability.