Authors
Isa Tuncay Batuk, Irem Karakuluk-Celebi
Published in
International journal of medical informatics. Volume 220. Pages 106633. Aug 02, 2026. Epub Aug 02, 2026.
Abstract
Hearing aid users frequently require accessible and immediate assistance for daily device management. This study aims to evaluate and compare the performance of two prominent Large Language Models (LLMs)-ChatGPT and Gemini, selected for their widespread public accessibility and market dominance-in providing accurate, comprehensible, and repeatable answers to frequently asked questions regarding hearing aids.
A comprehensive set of 44 user queries was divided into seven core categories. Responses generated by ChatGPT and Gemini were evaluated for comprehensibility and medical accuracy by three expert audiologists, utilizing official manufacturer manuals as a definitive gold standard. Inter-rater reliability was measured using the Intraclass Correlation Coefficient (ICC). Following the establishment of high consensus, repeatability was mathematically measured to objectively assess output similarity and eliminate human bias using a Natural Language Processing (NLP) algorithm (Cosine Similarity). Data were reported as medians and interquartile ranges (IQR), and pairwise statistical comparisons were conducted using the non-parametric Wilcoxon matched-pairs signed-rank test.
Inter-rater reliability was strong for both subjective domains (Accuracy ICC = 0.74; Clarity ICC = 0.74, p < 0.001). Both models demonstrated excellent performance in clarity (Median: 5.00, IQR: 0.00), with no statistically significant differences observed across any categories (p > 0.05). Regarding medical and technical accuracy, rank-based analyses revealed that Gemini showed significantly superior overall performance (p < 0.001), specifically excelling in the "Pairing" (p = 0.03) and "Before Using a Hearing Aid" (p = 0.03) categories. In the objective assessment of repeatability via NLP algorithms, both models exhibited high semantic consistency (ChatGPT Median: 0.91; Gemini Median: 0.93), with no statistically significant differences observed between them in any category (overall p = 0.30).
Both ChatGPT and Gemini generate highly comprehensible information for hearing aid users (Clarity Median: 5.00). However, their performance fluctuates depending on task complexity; while Gemini offers superior technical accuracy (p < 0.001), both maintain high but imperfect repeatability (Cosine Similarity > 0.90). Because neither model demonstrates complete diagnostic stability or perfect reproducibility across all clinical domains, they cannot currently be recommended as independent digital assistants. While LLMs show great promise as supplementary tools for patient education, professional audiological supervision remains essential to verify clinical accuracy and ensure patient safety.
PMID:
42571769
Bibliographic data and abstract were imported from PubMed on 10 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 5
- Comments 0