Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Comparative evaluation of large language models in interpreting the scientific literature on intraoral scanners across varying input levels.

Created on 09 Sep 2026

Authors

Pranay Jain, Fathima Banu R, Anand Kumar V, Sarvesh Kulkarni, Maheshwaran Ks

Published in

The Journal of prosthetic dentistry. Sep 08, 2026. Epub Sep 08, 2026.

Abstract

Dental professionals increasingly seek artificial intelligence (AI)-generated sources for concise information on emerging topics in advancement and scientific literature; however, concerns persist regarding the accuracy of AI tools and the reliability of responses generated based on the complexity of the prompt.
The purpose of this study was to evaluate and compare the response accuracy of 2 advanced large language models (LLMs): OpenAI ChatGPT-5.2 (Auto) and Google Gemini-3.0 Flash in interpreting and analyzing published data on intraoral scanners.
A total of 20 articles were selected through a structured PubMed search. Two independent senior prosthodontists were blinded and assessed every response generated by the LLMs at each prompt level. Prompts were fed at different levels (1-4) of complexity: Level 1 consisted of simplified prompts requiring high logical reasoning by the LLM, while Level 4 consisted of a structured prompt including methodology. The responses were evaluated across the 4 domains of relevance, scientific accuracy, clarity, and closeness to article content using a 5-point Likert scale. The nonparametric Friedman test and Wilcoxon signed-rank test with Bonferroni correction were conducted for statistical analysis (α=.05).
The median output for ChatGPT-5.2 for relevance of answer and clarity were higher than for Gemini-3.0 at the simpler prompt levels of Level 1 and Level 2 (P<.05), whereas the trend reversed at higher levels, with Gemini-3.0 performing significantly better at Level 3 (P<.001) and Level 4 (P<.001). In terms of scientific accuracy, ChatGPT-5.2 consistently achieved higher scores at Level 1 (P=.003) and Level 4 (P=.023), whereas Gemini-3.0 plateaued across levels (P>.05). In contrast, Gemini-3.0 outperformed ChatGPT-5.2 in closeness to article at Level 2 (P<.001) and 4 (P=.019).
The Gemini-3.0 Flash responses excelled in highly structured, direct-response questions to relevance, clarity, and closeness to the article, whereas the ChatGPT-5.2 response demonstrated greater scientific accuracy at all levels and performed better with simplified prompts that require high logical reasoning. Framing appropriate prompt input is essential for achieving better scientific accuracy when using AI tools for research analysis.

PMID:
42711193
Bibliographic data and abstract were imported from PubMed on 09 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 13
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement