Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Generative Artificial Intelligence Performance on University-Level Human Anatomy Examinations: A Structured Narrative Review and Proposed MATRIX-Anatomy Framework.

Created on 27 Aug 2026

Authors

Juan A Sanchis-Gimeno, Juan José Valenzuela-Fuenzalida, Alejandro Bruna-Mejías, Mathias Orellana-Donoso, Glen J Paton, Shahed Nalla

Published in

Clinical anatomy (New York, N.Y.). Aug 26, 2026. Epub Aug 26, 2026.

Abstract

Generative artificial intelligence (GenAI) can perform strongly on written anatomy examinations, but whether such scores represent anatomical competence remains uncertain because results vary with the model, assessment, protocol, modality, scoring, and comparator. We synthesized studies evaluating GenAI as the examinee in university-level human anatomy assessments and proposed MATRIX-Anatomy, a reporting and interpretive framework not yet externally validated. PubMed, Scopus, and Web of Science Core Collection were searched for records published from January 2022 to 11 July 2026. Eligible studies used university examinations, course item banks, or curriculum-aligned undergraduate benchmarks and reported quantitative performance. Three reviewers completed study selection, data extraction, and narrative synthesis by consensus. Of 115 records, 51 duplicates were removed, 64 were screened, and 15 studies were included. Leading systems scored 76% to 98% on text-based multiple-choice assessments. On a fixed 120-item set, accuracy increased from 45.8% with ChatGPT-3.5 to 86.7% with ChatGPT-5. Human comparisons were mixed. Visuospatial performance was weaker: ChatGPT-4o identified 22.26% of cadaveric structures after up to three attempts; ChatGPT-4.0 achieved 17.3% end-to-end accuracy on image-based anatomy; and ChatGPT-5.1 reached 74.4% on a surgical-anatomy subset. Repeated runs revealed volatility and consistently incorrect responses. The findings support supervised formative use with authoritative verification and retention of secure supervised, oral, constructed-response, visuospatial, and practical assessments. MATRIX-Anatomy specifies six domains (Model, Assessment, Testing protocol, Reference standard, Input, and eXternal validity) for reproducible reporting and defensible interpretation, but requires formal external validation.

PMID:
42655917
Bibliographic data and abstract were imported from PubMed on 27 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 11
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement