Authors
Hongxia Lu, Yan Huo, Ruisi Xie, Zhengyuan Qu, Yutong Li, Shengjin Wang, Haohan Zou, Yan Wang
Published in
Journal of ophthalmology. Volume 2026. Pages 4163980. Epub Aug 29, 2026.
Abstract
To evaluate the guideline knowledge alignment of two large language models (LLMs), GPT-5.5 Instant and DeepSeek-V4, and to determine their preclinical reliability as reference tools in refractive surgery.
Using the 38 evidence-based recommendations of the international keratorefractive lenticule extraction (KLEx) guidelines as the gold standard, both LLMs were evaluated in their default configurations. Two ophthalmologists independently assessed the clinical safety and medical accuracy of the model responses using a 5-point Likert scale (1-5 points). Agreement between each model's recommendation strength and the guideline was quantified by intraclass correlation coefficient (ICC), structural reliability by the DISCERN scale, and readability by the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL) indices.
Likert ratings did not differ between GPT-5.5 Instant (4.96 ± 0.206) and DeepSeek-V4 (4.89 ± 0.385; p = 0.134). Across 114 independent generations, the ICC for agreement with the guideline was 0.884 (95% CI, 0.836-0.918) for GPT-5.5 Instant and 0.739 (0.640-0.813) for DeepSeek-V4 (both p < 0.001). DISCERN scores were 70.18 ± 5.16 and 68.05 ± 5.41 (p = 0.085); FRE, 9.93 ± 8.33 and 3.74 ± 5.24 (p < 0.001); and FKGL, 16.07 ± 2.21 and 18.95 ± 1.96 (p < 0.001).
Both models aligned closely with the guidelines, with significant concordance in GRADE-based recommendation strength. They may serve as preclinical reference tools in refractive surgery but not as a substitute for specialist judgment.
PMID:
42668902
Bibliographic data and abstract were imported from PubMed on 30 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 6
- Comments 0