As conversational artificial intelligence becomes a routine tool for personal health queries, a newly published study warns that relying on AI chatbots for guidance on late-life depression carries significant clinical risks. Published in Frontiers in Psychiatry, the blinded benchmark study found that while popular models often answer baseline questions well, they frequently fail when confronting high-risk scenarios and crisis situations in older adults.
A Growing Reliance on AI for Mental Health
Patients and family caregivers are increasingly turning to generative AI chatbots for advice on depression and mental health crises. However, depression in older adults is rarely straightforward. The condition is uniquely complicated by factors such as cognitive decline, co-occurring physical illnesses, and polypharmacy—the concurrent use of multiple prescription medications.
According to the researchers, issues like physical frailty, self-neglect, and caregiver dependence make safe mental health guidance difficult for algorithms to navigate. When an automated system fails to account for these intersecting vulnerabilities, even broadly accurate medical facts can lead to dangerous outcomes.
Testing the Models: ChatGPT, Gemini, and Doubao
To evaluate how modern chatbots handle these complex scenarios, a research team led by Wei Xiao evaluated 90 clinically realistic questions across six domains of geriatric psychiatry. The prompts were divided equally into low-, moderate-, and high-risk strata to reflect real-world inquiries from patients and caregivers.
The researchers tested three prominent AI models: ChatGPT (GPT-5.5 Instant), Gemini (Gemini 3.5 Flash), and Doubao (Seed2.0 Pro), generating 270 primary-round responses. Two psychiatrists independently reviewed the anonymized outputs against prespecified clinical standards, with a third senior psychiatrist resolving any disagreements. A stratified subset of 30 questions was also resubmitted in separate sessions to test response consistency, totaling 360 evaluations.
The Findings: Acceptable on Basics, Alarmingly Weak on High Risk
The benchmark showed significant disparities in safety and accuracy. Overall, ChatGPT generated clinically acceptable responses in 78.9% of cases, Gemini in 72.2%, and Doubao in 60.0%. However, when evaluated specifically for complete geriatric appropriateness, performance fell across the board to 64.4%, 55.6%, and 43.3%, respectively.
Worryingly, major safety errors emerged in 5.6% of ChatGPT responses, 10.0% of Gemini responses, and 16.7% of Doubao responses. The models struggled most acutely when handling high-risk scenarios, where clinical acceptability plunged to 63.3% for ChatGPT, 50.0% for Gemini, and 36.7% for Doubao. In these critical situations, the bots frequently engaged in under-triage—failing to recommend urgent professional intervention in 16.7% of high-risk cases for ChatGPT, 26.7% for Gemini, and 36.7% for Doubao.
The Takeaway: Educational Aid, Not a Clinical Triage Tool
While the AI tools performed reasonably well on low-risk educational queries, the study identified critical vulnerabilities in crisis response. The models repeatedly missed indirect warning signs, provided incomplete safety guidance, and displayed fluctuating consistency on repeated questions, where test-retest agreement ranged from 73.3% to 90.0%.
The authors concluded that while general-purpose AI may support basic, low-risk educational inquiries, it must not be trusted to independently manage suicide risk, medication changes, or emergency triage in older adults. For complex geriatric care, human medical judgment remains indispensable.
Source
- Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage
- Researchers: Wei Xiao, Huanyu Zhang, Xiaoyi Chen, Jun Cai, 駱旭琛, Jiehua Deng
- Journal: Frontiers in Psychiatry
- Read the original study