Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Clarity Without Credibility? Human Versus AI Abstracts in Otolaryngology.

Created on 25 Jul 2026

Authors

Sholem Hack, Rebecca Attal, Lirit Levi, David T Liu, Mahdokht Manavi, Luca Giovanni Locatello, Maxime Fieux, Evan M de Joya, Joel M Gutovitz, Shahar Eliahu, Tal Hefetz, Tomer Boldes, Ameen Biadsee

Published in

World journal of otorhinolaryngology - head and neck surgery. Mar 23, 2026. Epub Mar 23, 2026.

Abstract

This study evaluated whether otolaryngologists can distinguish between human- and machine-written abstracts. The primary question was whether large language models (LLMs) produce abstracts comparable in clarity and usefulness to human-authored work, and whether reviewers can identify authorship with accuracy.
A blinded cross-sectional design was used. Forty-eight abstracts were evaluated, consisting of twenty-four human-authored abstracts and 24 generated by four LLMs. Human abstracts were drawn from articles published after July 2025 to minimize overlap with LLM training data. Twenty otolaryngologists independently reviewed all abstracts. Using a structured rubric, raters classified authorship, rated clarity, usefulness, and confidence on 5-point scales, and provided optional free-text explanations. Group comparisons were performed using chi-square and Mann-Whitney tests, with Kruskal-Wallis tests for model-level analyses.
Overall recognition accuracy was 44.7%. Human-written abstracts were more often misclassified as AI than AI-generated abstracts were mistaken for human. Human abstracts received significantly higher clarity and usefulness scores than LLM abstracts, though effect sizes were small. Confidence did not correlate with correctness, indicating miscalibration of rater judgments. Model-level performance varied. Grok-generated abstracts were most easily identified as AI, whereas GPT-5 and Claude 3.5 more frequently resembled human writing. Free-text rationales commonly referenced style, vagueness, or lack of detail when AI authorship was suspected.
LLMs generate abstracts that increasingly resemble human scientific writing, yet still lag in perceived usefulness and credibility. Clinicians were only moderately successful at detecting authorship and were frequently confident in incorrect classifications. These findings highlight both the promise and risks of AI-assisted scientific communication.

PMID:
42500058
Bibliographic data and abstract were imported from PubMed on 25 Jul 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 1
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement