Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Large language model accuracy in inhaler technique counselling for asthma and COPD: A comparison of free and paid models across ten devices.

Created on 17 Aug 2026

Authors

Abdurrahman Koç, Ferhat Sağun, Necmettin Öğe, Sami Avcil, Ali Tolga Çelik, Muhammet Ali Takeş, Bekir Sunay

Published in

Digital health. Volume 12. Pages 20552076261478506. Epub Aug 14, 2026.

Abstract

To compare ten large language models (LLMs) from seven AI companies, across free and paid tiers, on inhaler technique instruction accuracy for ten devices, benchmarked against GINA 2025 and GOLD 2026 strategy reports.
In a cross-sectional, blinded evaluation, ten LLMs (ChatGPT Free/Plus, Claude Free/Pro, Gemini, Google AI Pro, DeepSeek, Microsoft Copilot, Perplexity AI, Meta AI) were queried between 10-17 March 2026 via each vendor's official web interface. Fifty standardised prompts (10 devices × 5 question types) were submitted in triplicate on different days, yielding 1500 outputs. Two pulmonologists and one allergist scored seven metrics step-completion rate, critical and non-critical error counts, step-sequencing accuracy, safety-warning score, and Likert-scaled overall accuracy and patient comprehensibility against a gold standard derived from GINA 2025, GOLD 2026, the ERS/ISAM Task Force consensus, and manufacturer leaflets.
Paid tiers outperformed free tiers on five of seven metrics (Mann-Whitney U; all p < 0.001), with a 24-fold reduction in critical errors (0.48 → 0.02) and a 12.9-point gain in safety-warning coverage (62.4% → 75.3%). Google AI Pro achieved the highest overall accuracy (4.98 ± 0.13); Microsoft Copilot (3.23) and Perplexity AI (2.83) ranked lowest. Claude Free surpassed Claude Pro on safety warnings (99.2% vs. 66.7%; r = 0.86). Test-retest reliability met ICC ≥ 0.75 for only two metrics.
Paid subscriptions yield meaningful but non-uniform gains in inhaler-instruction accuracy. Safety-warning omissions and test-retest variability argue against autonomous LLM use; these tools are best deployed as adjuncts to clinician- and pharmacist-led teach-back education.

PMID:
42605352
Bibliographic data and abstract were imported from PubMed on 17 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 6
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement