Life sciences · Journal article
npj Digital Medicine · July 7, 2026
Early or partial results. Treat as a signal, not a conclusion.
This proof-of-concept study compares ten large language models to 108 early-career clinicians on structured psychiatric assessment across 100 items using simulated interviews. GPT-5.1 and Gemini-3-Pro-Preview achieved 0.72 accuracy (64th percentile of clinician distribution), with variable performance by diagnostic scenario (depression 0.81, mania 0.76, schizophrenia 0.60). The findings are preliminary and require validation in real patient interviews and prospective study designs before clinical utility can be established.
Proof-of-concept benchmarking study with simulated interviews and expert consensus reference. 108 early-career clinicians; three simulated psychiatric interviews; no real patients studied. Intervention: Assessment by 10 large language models (GPT-5.1, Gemini-3-Pro-Preview, and 8 others) on AMDP 100-item structured psychiatric assessment using majority voting and AMDP definitions as context. Compared with: Assessment by 108 early-career clinicians viewing audiovisual recordings; expert consensus panel as reference standard. n = 108.
GPT-5.1 and Gemini-3-Pro-Preview achieved highest accuracy of 0.72 using majority voting with AMDP definitions as context GPT-5.1 per-scenario accuracies: depression 0.81 vs clinician mean 0.79, mania 0.76 vs 0.68, schizophrenia 0.60 vs 0.58 Clinicians and LLMs showed distinct error profiles: clinicians over-inferred symptom presence; LLMs marked items as 'not assessable' more frequently (19.4% vs 11.4%, p < 0.001)
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
These findings suggest LLMs may assist in structured psychiatric assessment, but the use of simulated interviews, expert consensus reference, and early-career clinician raters limits immediate clinical applicability. Real-world validation and prospective study designs in actual patient encounters are required before any clinical integration can be recommended.
Proof-of-concept study using simulated interviews and expert consensus reference rather than clinical outcomes; demonstrates feasibility but requires validation in real patient settings before clinical application.
As stated by the source record.
Quoted from the source exactly as published.
These findings suggest LLMs may assist in structured psychiatric assessment, but the use of simulated interviews, expert consensus reference, and early-career clinician raters limits immediate clinical applicability. Real-world validation and prospective study designs in actual patient encounters are required before any clinical integration can be recommended.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Psychiatry's reliance on language makes LLMs a natural tool for psychopathological assessment, yet structured, item-level assessments from psychiatric clinical interviews remain under-researched. In this proof-of-concept study, 10 LLMs assessed transcripts of three simulated psychiatric interviews across all 100 items of the Association for Methodology and Documentation in Psychiatry (AMDP) system, benchmarked against 108 early-career clinicians rating full audiovisual recordings, using an expert consensus panel as reference. GPT-5.1 and Gemini-3-Pro-Preview achieved the highest accuracy (0.72; 64th percentile of the clinician distribution) using majority voting across three runs with AMDP definitions as context. GPT-5.1, selected for a marginal advantage, showed per-scenario accuracies of 0.81 (depression), 0.76 (mania), and 0.60 (schizophrenia) versus clinician means of 0.79, 0.68, and 0.58. Clinicians and LLMs showed distinct error profiles: clinicians tended to over-infer symptom presence, whereas LLMs more conservatively flagged items as "not assessable" - most pronounced for observation-dependent items but present even for text-assessable items (19.4% vs. 11.4%, p < 0.001). In post hoc simulated disagreement resolutions (2091 clinician pairs; 35.5% disagreements), LLM and board-certified supervision were associated with more accurate resolutions than unsupervised random clinician selection (p < 0.0002). These proof-of-concept findings require validation in real patient interviews, larger samples, and prospective studies integrating multimodal input.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.