Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Benchmarking five large language models in medical genetics: a bilingual comparative evaluation using published and novel expert-authored questions.

Created on 30 Sep 2026

Authors

Özge Beyza Gündoğdu Öğütlü, Benjamin D Solomon, Yusuf Selman Çelik

Published in

Frontiers in medicine. Volume 13. Pages 1938886. Epub Sep 15, 2026.

Abstract

This study asked whether five contemporary large language models answer medical genetics multiple-choice questions with equivalent accuracy on published versus novel items and across English and Turkish, and sought to characterize the errors that persist. Five models (GPT-5.2, Gemini 3 Pro, Claude Sonnet 4.6, Grok 4, and DeepSeek-V3.2) answered 100 four-option questions (50 from a published board review; 50 novel, expert-authored items absent from any database) in English and Turkish, yielding 1,000 responses. Correctness was modeled with item-clustered generalized estimating equation and Bayesian mixed-effects logistic regression (the latter as a prespecified sensitivity analysis); question provenance and language were tested for equivalence (item-clustered two one-sided tests, ±5-percentage-point margin), and the paired language effect with the McNemar test. Inter-model agreement and error concordance were examined. Overall accuracy was 97.7%. In the item-clustered GEE, only Gemini 3 Pro nominally exceeded the lowest-performing model; this imprecise contrast did not remain significant after Holm correction for the four secondary model comparisons, whereas the Bayesian sensitivity analysis additionally yielded a credible interval excluding 1 for GPT-5.2 versus Grok 4. Accuracy was statistically equivalent within the prespecified margin for published versus novel items (difference, -1.4 percentage points; item-clustered 90% CI, -4.1 to +1.3; item-clustered TOST P = 0.014) and for English versus Turkish (difference, -0.6 points; paired item-level TOST P < 0.001; McNemar P = 0.58). Inter-model agreement on the exact option chosen was high (Fleiss κ, 0.94-0.97). All five items with two or more errors failed concordantly, every erring model selecting the same wrong option. Beyond high accuracy, distinct model families converged on identical answers and, on the hardest items, on identical errors, a pattern consistent with shared learned associations and/or overlapping training data, although this behavioral comparison cannot identify the underlying mechanism. In an exploratory analysis, residual errors showed concentrated and concordant patterns, clustering in multistep Bayesian reasoning and evolving facts, suggesting that a second model may provide limited independent protection on such items. Performance on this restricted multiple-choice benchmark does not establish readiness for clinical genetic counseling, variant interpretation, or quantitative risk assessment, and human oversight remains essential.

PMID:
42812270
Bibliographic data and abstract were imported from PubMed on 30 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 1
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement