Authors
Yi-Lin Wang, Liao-Lin Chen, Ping-Ping Sun, Zhen-Zhen Ma, Ye-Ping Chen, Yang-Shuo Ge, Xin-Hui Huang, Chun-Meng Huang, Jia-Wei Du, Ting-Ting Meng, Dao-Fang Ding
Published in
Digital health. Volume 12. Pages 20552076261484840. Epub Sep 05, 2026.
Abstract
To evaluate large language models (LLMs) for stroke care in guideline-based question answering (Q&A) and individualized draft rehabilitation-plan generation, with explicit assessment of response quality, readability, and repeated-generation score stability.
A two-stage study was conducted. Stage 1: Four LLMs generated best answers to 100 stroke-related questions; the top two models by accuracy advanced to Stage 2. In Stage 2a, ChatGPT-o3 and DeepSeek answered 20 guideline-derived clinical questions, each repeated three times. In Stage 2b, the two models generated individualized rehabilitation plans for 60 de-identified stroke cases, with three independent generations per case. Three senior clinicians evaluated outputs across correctness, completeness, readability, helpfulness, and safety using 5-point Likert scales. Chinese readability was assessed using the LDU-TGP platform, and English readability was assessed using the Flesch-Kincaid Grade Level. Generalized estimating equations were used for Stage 2a and Stage 2b comparisons to account for repeated generations clustered within clinical questions and patient cases, respectively.
In Stage 1, DeepSeek achieved the highest accuracy (91%), followed by ChatGPT-o3 (90%), Gemini (85%), and ChatGPT-4o (83%). In Stage 2a, no between-model differences across the five clinician-rated domains remained statistically significant after correction for multiple comparisons. ChatGPT-o3 generated clinical Q&A responses with a higher recommended reading age and English Flesch-Kincaid Grade Level, whereas the difference in Chinese Reading Difficulty Score did not remain significant after correction. In Stage 2b, ChatGPT-o3 significantly outperformed DeepSeek across correctness, completeness, readability, helpfulness, and safety in the GEE analysis. Repeated-generation score stability did not differ significantly between models.
Within this single-center, expert-rated benchmark, ChatGPT-o3 and DeepSeek performed comparably across most guideline-based Q&A dimensions, whereas ChatGPT-o3 achieved higher clinician-rated scores for draft rehabilitation-plan generation. These findings do not establish clinical effectiveness. Prospective patient-centered validation and clinician-supervised review are required before implementation in routine stroke rehabilitation.
PMID:
42703353
Bibliographic data and abstract were imported from PubMed on 07 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 10
- Comments 0