Authors
Öznur Küçük Keleş, Zeynep Betül Arslan
Published in
Odontology. Aug 10, 2026. Epub Aug 10, 2026.
Abstract
This study aims to compare the diagnostic accuracy, appropriateness of treatment planning, and source citation performance of five large language models ChatGPT-4o (Free), ChatGPT-5.1 Plus, Microsoft Copilot, Google Gemini, and DeepSeek-R1 in root resorption scenarios. In December 2025, twelve clinical scenarios were created based on the classification of the European Society of Endodontology and each scenario was presented to all chatbots over four consecutive days. All responses were evaluated using a blinded assessment protocol and a binary scoring system. A total of 720 observations (12 cases × 4 repetitions × 3 criteria per model) were analyzed. The collected data were analyzed using chi-square, Fisher's exact, and Cochran Q tests. In terms of diagnostic accuracy, Microsoft Copilot (79.2%), ChatGPT-5.1 (77.1%), and ChatGPT-4o (Free) (75%) showed the highest performance. Google Gemini (68.8%) demonstrated a moderate level of accuracy, while DeepSeek (39.6%) showed markedly low performance. All models exhibited high accuracy in treatment plan recommendations, and no statistically significant differences were detected. Regarding citation accuracy, Copilot ranked first with 100% accuracy. Although large language models present potential as supportive decision-making tools in the evaluation of root resorption, diagnostic inconsistencies, limitations in source accuracy, and variability in responses restrict their independent use in clinical applications. Therefore, the outputs generated by these models should be interpreted cautiously within clinical decision-making processes.
PMID:
42573919
Bibliographic data and abstract were imported from PubMed on 10 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 2
- Comments 0