Life sciences · Journal article
Healthcare · October 6, 2026
No summary has been generated for this record yet. What follows is drawn from its source metadata only.
Journal article.
No findings were extractable from the material analysed.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
This record has not been graded across any dimension yet. Treat the label above as provisional and read the source.
What is missing. This record has no bottom line, key findings, reported figures, evidence dimensions. That is a gap in the analysis, not a judgement about the study.
Background/Objectives: Artificial intelligence (AI) chatbots are increasingly used for healthcare information retrieval, education, and clinical support; however, concerns remain regarding the accuracy and completeness of AI-generated information in preventive medicine, where evidence-based recommendations, epidemiological interpretation, risk assessment, and public health guidance require reliable responses. Comparative evidence across multiple preventive medicine domains remains limited, particularly using standardized expert-validated reference answers. This study aimed to evaluate and compare the accuracy and completeness of responses generated by five AI chatbots across major preventive medicine domains. Methods: A comparative cross-sectional benchmarking study evaluated ChatGPT, DeepSeek, Perplexity AI, Claude AI, and Gemini Advanced using 15 expert-validated questions covering General Preventive Medicine, Clinical Preventive Medicine, Epidemiology and Biostatistics, Occupational Medicine, and Social and Behavioral Science. Responses were independently evaluated by two investigators using predefined accuracy and completeness scoring frameworks. Between-model comparisons were conducted using mixed-effects models accounting for clustering by question, and inter-rater reliability was assessed using agreement statistics. Results: Claude AI achieved the highest overall accuracy (5.67 ± 0.41) and completeness (2.87 ± 0.35), followed by ChatGPT (accuracy, 5.47 ± 0.66; completeness, 2.67 ± 0.49), whereas Gemini Advanced demonstrated the lowest overall performance. Epidemiology and Biostatistics showed the highest pooled completeness score (2.60; 95% CI: 2.32–2.88). Mixed-effects analysis accounting for clustering by question demonstrated a significant overall difference in accuracy among chatbot models (p < 0.001), while inter-rater agreement was high, supporting the consistency of the scoring procedure. Conclusions: Substantial variability in accuracy and completeness was observed among the evaluated AI chatbots, with Claude AI and ChatGPT demonstrating comparatively stronger performance within this specific question set. Given the limited number of unique questions and the continuously evolving nature of AI models, these findings should be considered exploratory and should not be generalized as definitive evidence of model superiority or clinical utility. Larger studies using more diverse question sets and standardized evaluation frameworks are required to confirm these findings.