Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Temporal linguistic shifts in oncology randomized controlled trials following large language model availability: A corpus analysis of 21,392 publications.

Created on 29 Jul 2026

Authors

Amina Silva, Vanessa Silva E Silva

Published in

European journal of cancer (Oxford, England : 1990). Volume 245. Pages 116963. Jul 24, 2026. Epub Jul 24, 2026.

Abstract

Public availability of large language models (LLMs) from late 2022 has raised concerns about AI-assisted writing in scientific publishing. Oncology randomized controlled trials (RCTs) underpin cancer treatment guidelines and regulatory decisions worldwide, yet whether linguistic changes have followed the emergence of LLMs has not been examined at scale.
We conducted a retrospective corpus linguistics analysis of 21,392 oncology RCTs indexed in PubMed (2019-2026), stratified into pre-LLM (2019-2022; n = 11,308) and post-LLM (2023-2026; n = 10,084) eras. Full text was retrieved for 10,483 papers (49.0%) via PubMed Central; abstracts were used otherwise. The primary outcome was the frequency of 34 formulaic phrases documented as disproportionately prevalent in post-LLM biomedical text. A secondary outcome applied the GRIM test. Mann-Whitney U tests, chi-squared analyses, and multivariable linear regression.
Post-LLM papers contained significantly more formulaic phrases than pre-LLM papers (mean 0.96 [SD 1.50] vs. 0.65 [SD 1.14]; U=51,334,465, Z = -14.56, p < 0.001; r = 0.10). The increase was consistent across full-text and abstract-only subgroups and largest in discussion sections (1.23 vs. 0.88; +40%). Year-by-year analysis showed a progressive rise from 2019 (0.45) through 2025 (1.13). GRIM results were limited by data availability (48 papers; null result, p = 0.085).
This study identified a statistically significant increase in formulaic linguistic patterns in oncology RCTs following the public release of LLMs. Although these population-level findings cannot confirm AI-assisted writing in individual papers or authors, they are consistent with increasing LLM use and highlight the need for transparent AI disclosure, ongoing surveillance of scientific writing, and evidence-informed editorial policies.

PMID:
42520591
Bibliographic data and abstract were imported from PubMed on 29 Jul 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 8
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement