Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Beyond Accuracy: A Mixed-Methods Audit of Chain-of-Thought Failures in LLM-Based COVID-19 Vaccine Stance Detection.

Created on 09 Aug 2026

Authors

Andreas Praschk, Valentin Fischill-Neudeck, Thomas Caspari, Hans-Peter Wiesinger

Published in

Journal of medical systems. Volume 50. Issue 1. Aug 08, 2026. Epub Aug 08, 2026.

Abstract

This mixed-methods study assessed whether reasoning-enabled large language models (LLMs) can classify stances towards COVID-19 vaccination on X (formerly Twitter) and whether model-generated chain-of-thought (CoT) summaries contain reasoning failures relevant to transparent and auditable public health applications. Zero-shot stance classification by o4-mini and Gemini 2.5 Flash (Gemini) was evaluated on 3,060 rehydrated COVID-19 vaccination tweets against human-annotated labels (positive, negative, neutral). We reported accuracy and macro-F1, measured CoT availability, and qualitatively analysed dual-error cases (tweets misclassified by both models) using Mayring's content analysis guided by the FUTURE-AI framework. At each model's best-performing setting, both models reached macro-F1 around 0.8, with o4-mini outperforming Gemini (accuracy 0.819 vs. 0.799, McNemar p = 0.0015; Δmacro-F1 = 0.020, 95% CI 0.008-0.032). Under the reasoning-intensive settings, CoT availability differed: Gemini returned a reasoning summary for all tweets, whereas o4-mini did so for 64.7%. Among 1,981 tweets with CoTs from both models, 295 (14.9%) were dual-errors; in 88.8%, both models produced the same wrong label, suggesting shared failure modes. Qualitatively, both models showed the same errors: target confusion (policy vs. vaccine), literal readings of sarcasm, and label-rationale mismatches, recurring across models despite their markedly different CoT lengths. Reasoning LLMs can therefore classify stance accurately, but their readiness for transparent public health applications depends on whether a CoT is available at all and whether it is coherent with the label it accompanies (label-rationale coherence). CoT availability, label-rationale coherence, and safeguards against systematic reasoning failures offer candidate explainability-readiness metrics, alongside accuracy, for trustworthy digital epidemiology.

PMID:
42570145
Bibliographic data and abstract were imported from PubMed on 09 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 6
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement