Machine Learning in Healthcare · Journal article
Frontiers in Psychiatry · September 11, 2026
Encouraging direction, but not yet definitive.
In a blinded benchmarking of three large language models answering 90 questions about late-life depression, ChatGPT achieved 78.9% clinically acceptable responses, Gemini 72.2%, and Doubao 60.0%, with major safety errors in 5.6%, 10.0%, and 16.7% respectively. Performance degraded markedly on high-risk questions (clinically acceptable responses 63.3%, 50.0%, 36.7%), with under-triage rates of 16.7%, 26.7%, and 36.7%. The findings support use of these models only for low-risk educational purposes and explicitly contraindicate their use for emergency triage, suicide-risk management, or medication guidance in older adults.
Blinded, paired benchmarking study with independent dual evaluation and adjudication. Benchmark question set (n=90) designed to reflect patient and caregiver inquiries about late-life depression, covering six geriatric psychiatry domains: cognitive change, multimorbidity, frailty, polypharmacy, self-neglect, and suicide risk. Questions stratified equally across low-risk, moderate-ri…. Intervention: Three large language models: ChatGPT (GPT-4.5 Instant via ChatGPT), Gemini (Gemini 3.5 Flash via Gemini), and Doubao (Seed2.0 Pro via Doubao).. Compared with: Prespecified item-specific reference standards reflecting clinical accuracy, safety, geriatric appropriateness, warning-sign recognition, and triage criteria for late-life depression..
Clinically acceptable responses: ChatGPT 78.9%, Gemini 72.2%, Doubao 60.0% (overall P<0.001) Major safety errors: ChatGPT 5.6%, Gemini 10.0%, Doubao 16.7% (FDR-adjusted q=0.023) Complete geriatric-specific appropriateness: ChatGPT 64.4%, Gemini 55.6%, Doubao 43.3%
Major safety errors: ChatGPT 5.6%, Gemini 10.0%, Doubao 16.7% (FDR-adjusted q=0.023)
Clinicians and patients should recognize that while general-purpose LLMs can provide accurate educational information on low-risk depression topics, they demonstrate significant and clinically unacceptable failure rates on high-risk scenarios including suicide assessment, medication guidance, and emergency triage. These models should not be used independently for safety-critical decisions in older adults with depression, though they may supplement low-risk information delivery after clinical assessment.
A rigorous blinded benchmarking study of three LLMs on a clinically relevant task (late-life depression triage) shows ChatGPT and Gemini perform moderately well on low-risk questions but all models fail substantially on high-risk scenarios, supporting cautious use only in low-risk educational contexts.
As stated by the source record.
Quoted from the source exactly as published.
Clinicians and patients should recognize that while general-purpose LLMs can provide accurate educational information on low-risk depression topics, they demonstrate significant and clinically unacceptable failure rates on high-risk scenarios including suicide assessment, medication guidance, and emergency triage. These models should not be used independently for safety-critical decisions in older adults with depression, though they may supplement low-risk information delivery after clinical assessment.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Background General-purpose large language models are increasingly used by patients and caregivers to obtain mental health information and guidance about when professional care is required. In late-life depression, broadly accurate information may nevertheless be unsafe when cognitive change, multimorbidity, frailty, polypharmacy, self-neglect, caregiver dependence, or suicide risk is not adequately recognized. This study compared the clinical accuracy, safety, geriatric-specific appropriateness, triage performance, and response consistency of three large language models when answering patient- and caregiver-centered questions about late-life depression. Methods We conducted a blinded, paired benchmarking study using 90 questions covering six geriatric psychiatry domains and equally distributed across low-, moderate-, and high-risk strata. Each question was submitted independently to GPT-5.5 Instant via ChatGPT, Gemini 3.5 Flash via Gemini, and Seed2.0 Pro via Doubao, generating 270 primary-round responses. A stratified subset of 30 questions was resubmitted in separate conversations to assess test-retest consistency, yielding 360 responses overall. Two psychiatrists independently evaluated anonymized outputs against prespecified item-specific reference standards, with clinically important disagreements adjudicated by a third senior psychiatrist. The primary outcome was the proportion of clinically acceptable responses, defined using accuracy, clinical safety, geriatric appropriateness, warning-sign recognition, and triage criteria. Results Clinically acceptable responses were generated for 78.9% of questions by ChatGPT, 72.2% by Gemini, and 60.0% by Doubao (overall P<0.001). The difference between ChatGPT and Gemini was not statistically significant, whereas both outperformed Doubao. This model ranking remained unchanged under alternative core-safety and more stringent optimal-response definitions. Complete geriatric-specific appropriateness was achieved in 64.4%, 55.6%, and 43.3% of responses, respectively. Major safety errors occurred in 5.6% of ChatGPT responses, 10.0% of Gemini responses, and 16.7% of Doubao responses (raw P = 0.015; FDR-adjusted q=0.023). Performance declined substantially with increasing clinical risk. In exploratory post hoc caregiver-centered analyses, clinical acceptability was 76.7% for ChatGPT, 70.0% for Gemini, and 53.3% for Doubao, while explicit caregiver-directed action was present in 80.0%, 70.0%, and 53.3% of responses, respectively. Among high-risk questions, clinically acceptable response rates were 63.3%, 50.0%, and 36.7%, while under-triage occurred in 16.7%, 26.7%, and 36.7% of responses, respectively. Test-retest clinical consistency was highest for ChatGPT (90.0%), followed by Gemini (83.3%) and Doubao (73.3%). Conclusions The three models answered many late-life depression questions accurately and safely, but none demonstrated consistently reliable performance across complex or high-risk scenarios. Clinically important weaknesses involved geriatric-specific interpretation, recognition of indirect risk signals, crisis-response completeness, under-triage, and response stability. General-purpose large language models may support selected low-risk educational tasks but should not independently guide emergency triage, medication changes, suicide-risk management, or other safety-critical decisions in older adults.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.