Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

A Quality Assessment Rubric for Artificial Intelligence-Generated Patient-Friendly Radiology Reports.

Created on 09 Sep 2026

Authors

Bonnie A Armstrong, Arogya Koirala, Hye Sun Na, Andrew Johnston, Marcello Chang, Donghe Lyu, Lina Cheuy, Zhongnan Fang, Akshay Chaudhari, David B Larson

Published in

AJR. American journal of roentgenology. Sep 09, 2026. Epub Sep 09, 2026.

Abstract

Background: Artificial intelligence (AI) tools are being used to translate radiology reports into plain language, but translation errors may compromise comprehension and safety. Objective: To develop and evaluate a rubric for assessing the quality and safety of AI-generated patient-friendly radiology reports. Methods: In this prospective study (conducted from February 2025 to December 2025), survey-workshop cycles, involving lay participants and a multidisciplinary panel, were used to develop a rubric for grading AI-generated patient-friendly report quality across core attributes and determining whether such reports are safe for patient distribution. ChatGPT-4.1 and Claude-4.0 were used to generate patient-friendly reports of varying quality based on radiology report impressions from a public dataset and prespecified quality targets across attributes. Research-team members, additional lay participants and radiologists, and ChatGPT-5 evaluated these patient-friendly reports using the rubric. Results: Development included 19 participants (39±7 years; 11 women, 8 men); evaluation included six research-team members and 111 additional participants (47±18 years; 43 women, 68 men). The final rubric included five core attributes-clarity, content, certainty, tone, verbosity-each graded on a 3-point scale; by the rubric's decision rule, patient-friendly reports assessed as grade 1 (unsafe or unacceptable) in any attribute other than verbosity are unsafe for distribution and warrant withholding. Lay and radiologist research-team members (n=3 participants each; 60 reports) had almost-perfect intergroup agreement (α=0.87) for overall grade assignments. Additional lay (n=19) and radiologist (n=12) participants, each evaluating six reports, had moderate (α=0.51) and substantial (α=0.65) interreader agreement, respectively, for overall grade assignments and 91.2% and 95.8% agreement, respectively, between subjective and rubric rule-based distribution decisions. In wider field testing, 80 lay participants (480 reports) had moderate agreement (κ=0.43) with prespecified reference-standard grades; subjective distribution decisions had 73.5% agreement with rule-based distribution decisions. Across 480 reports, AI had moderate agreement (κ=0.44) with prespecified reference-standard grades; rule-based distribution decisions using AI-assigned grades had 88.1% agreement with rule-based distribution decisions using reference-standard grades. Conclusions: The rubric may provide a standardized safeguard before release of AI-generated patient-friendly reports. Clinical Impact: Although requiring further training and validation, AI rubric application could enable scalable quality assurance and safer clinical integration of AI-generated communications.

PMID:
42714439
Bibliographic data and abstract were imported from PubMed on 09 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 21
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement