Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Human Judgment and the Limits of Artificial Intelligence for Automated Rank Order Lists in Diagnostic Radiology Residency Selection.

Created on 12 Sep 2026

Authors

Cody H Savage, Rydhwana Hossain, James Mac Tonascia, Athanasios Pavlou, Charles S Resnik, Florence X Doo, Elana B Smith

Published in

Journal of the American College of Radiology : JACR. Sep 11, 2026. Epub Sep 11, 2026.

Abstract

To evaluate whether large language models (LLMs) can reliably reproduce a residency program's rank order list (ROL) from pre-interview application data alone versus with human interviews included, and to determine whether automated ranking tools safely mirror human consensus or inadvertently introduce systematic displacements against specific applicant subgroups.
In this single-institution pilot study of 148 applicants interviewed during the 2025-2026 cycle, seven LLM configurations (three proprietary models [each at "medium" and "high" reasoning levels], and one open-weight model) each generated 10 independent ROLs from de-identified application data under two conditions: pre-interview application-only data alone, and then with interviewer scores added (140 lists total). Non-model baselines (sorting solely by interview score or USMLE Step 2 CK score) were also established. Lists were compared with the committee's final ROL using a truth-anchored weighted Kendall's τ and precision@k (k=10-40); rank displacement by applicant subgroup was also assessed.
Producing the committee's ROL required approximately 390 faculty-hours to order 148 applicants for 7 positions. In the application-only condition, all LLM configurations showed low agreement with the final ROL (median τ 0.15-0.36). In the positive control condition with human interviews included, agreement rose to excellent (median τ 0.84-0.93). Sorting applicants by summed interviewer score alone reproduced the final list at τ-b=0.83, whereas sorting by USMLE Step 2 CK score alone produced τ-b=0.12. Increasing reasoning level from "medium" to "high" did not consistently improve agreement. In the application-only condition, LLM rankings systematically placed female, international, and non-MD applicants below the human committee's list (q<0.05), while program signalers were placed higher.
LLMs failed to replicate the program's final ROL from pre-interview application-only data alone. The positive control including the human interview confirms this human-LLM lack of agreement may stem from the limited application-only data provided, rather than the models' inability to perform ranking tasks once provided the additional human judgment information. Importantly, these findings do not establish whether the LLM model or the human committee produced the "better" rank order list; this data suggests that LLM rankings generated solely from pre-interview data systematically differed from the human committee's consensus, and may disproportionately penalize certain subgroups. Our findings suggest that AI should not be used as a primary ROL generator or a replacement for human judgment, however may serve supportive roles in tasks such as data retrieval and retrospective auditing.

PMID:
42727622
Bibliographic data and abstract were imported from PubMed on 12 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 6
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement