Authors
Yosef Adiniaev, Mahmud Omar, Oved Daniel, Tohar M Timor, Yiftach Barash, Olga R Brook, Eyal Klang, Alon Gorenshtein
Published in
Neurological sciences : official journal of the Italian Neurological Society and of the Italian Society of Clinical Neurophysiology. Volume 47. Issue 8. Jul 13, 2026. Epub Jul 13, 2026.
Abstract
Dementia affects over 55 million people worldwide. Mild cognitive impairment (MCI) often precedes Alzheimer's disease (AD). Clinical management requires integrating uncertain evidence from neuropsychological testing, neuroimaging, and biomarkers. Large language models (LLMs) also generate probabilistic outputs, but whether they can reliably support diagnostic, therapeutic, or educational tasks in AD and MCI has not been systematically examined.
We searched PubMed, Scopus, and PubMed Central (January 2023 to April 2026) for studies evaluating generative LLMs on clinical tasks in Alzheimer's disease (AD) or mild cognitive impairment (MCI). Risk of bias was assessed using QUADAS-AI and AXIS. Narrative synthesis followed the SWiM guideline.
CRD420261372436.
Eleven studies were included: diagnosis (n = 3), treatment guidance (n = 2), and patient/caregiver education (n = 8); two studies contributed to multiple domains. Diagnostic models achieved high internal accuracy (0.94-0.97) but declined on external validation; three-way classification accuracy dropped approximately 7% points, and MMSE-prediction R² collapsed from 0.90 to 0.25 on an external dataset. Treatment guidance approached but did not match structured clinical guidelines. Educational outputs were rated moderate to high quality but lacked source attribution and exceeded recommended reading levels; retrieval augmentation improved usability without improving accuracy. Hallucination was quantified in only 2 of 11 studies, and no study evaluated prospective clinical use.
Current evidence does not support the use of LLMs for diagnosis, treatment selection, or patient education in AD/MCI without clinician oversight. These findings reflect the specific model versions, prompting strategies, and evaluation conditions in place at the time of each study, and are further limited by small heterogeneous evaluations, sparse hallucination measurement, and absence of prospective clinical validation.
PMID:
42440193
Bibliographic data and abstract were imported from PubMed on 13 Jul 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 11
- Comments 0