Generative Artificial Intelligence Performance onUniversity-Level Human Anatomy Examinations:A Structured Narrative Review and ProposedMATRIX-Anatomy Framework

dc.contributor.authorSanchis-Gimeno, Juan A.
dc.contributor.authorValenzuela-Fuenzalida, Juan José
dc.contributor.authorBruna-Mejías, Alejandro
dc.contributor.authorOrellana-Donoso, Mathias
dc.contributor.authorPaton, Glen J.
dc.contributor.authorNalla, Shahed
dc.coverage.spatialEstados Unidos
dc.date.accessioned2026-09-28T15:23:50Z
dc.date.available2026-09-28T15:23:50Z
dc.date.issued2026-08-20
dc.description.abstractGenerative artificial intelligence (GenAI) can perform strongly on written anatomy examinations, but whether such scores represent anatomical competence remains uncertain because results vary with the model, assessment, protocol, modality, scoring, and comparator. We synthesized studies evaluating GenAI as the examinee in university-level human anatomy assessments and proposed MATRIX-Anatomy, a reporting and interpretive framework not yet externally validated. PubMed, Scopus, and Web of Science Core Collection were searched for records published from January 2022 to 11 July 2026. Eligible studies used university examinations, course item banks, or curriculum-aligned undergraduate benchmarks and reported quantitative performance. Three reviewers completed study selection, data extraction, and narrative synthesis by consensus. Of 115 records, 51 duplicates were removed, 64 were screened, and 15 studies were included. Leading systems scored 76% to 98% on text-based multiple-choice assessments. On a fixed 120-item set, accuracy increased from 45.8% with ChatGPT-3.5 to 86.7% with ChatGPT-5. Human comparisons were mixed. Visuospatial performance was weaker: ChatGPT-4o identified 22.26% of cadaveric structures after up to three attempts; ChatGPT-4.0 achieved 17.3% end-to-end accuracy on image-based anatomy; and ChatGPT-5.1 reached 74.4% on a surgical-anatomy subset. Repeated runs revealed volatility and consistently incorrect responses. The findings support supervised formative use with authoritative verification and retention of secure supervised, oral, constructed-response, visuospatial, and practical assessments. MATRIX-Anatomy specifies six domains (Model, Assessment, Testing protocol, Reference standard, Input, and eXternal validity) for reproducible reporting and defensible interpretation, but requires formal external validation.
dc.identifier.citationClinical Anatomy (2026) pp. 1-15
dc.identifier.doihttps://doi.org/10.1002/ca.70218?utm_source=gemini
dc.identifier.issn0897-3806
dc.identifier.orcidhttps://orcid.org/0000-0002-1781-062X
dc.identifier.urihttps://hdl.handle.net/20.500.12254/7745
dc.language.isoen
dc.publisherWiley Periodicals LLC
dc.rightsAtribución-NoComercial-CompartirIgual 3.0 Chile (CC BY-NC-SA 3.0 CL)
dc.rights.urihttp://creativecommons.org/licenses/by-nc-sa/3.0/cl/
dc.subjectAnatomy education
dc.subjectAssessment validity
dc.subjectClinical anatomy
dc.subjectGenerative artificial intelligence
dc.subjectLarge language models
dc.subjectUniversity examinations
dc.titleGenerative Artificial Intelligence Performance onUniversity-Level Human Anatomy Examinations:A Structured Narrative Review and ProposedMATRIX-Anatomy Framework
dc.typeArticle
Archivos
Bloque original
Mostrando 1 - 1 de 1
No hay miniatura disponible
Nombre:
Paper 19 2026.pdf
Tamaño:
809.67 KB
Formato:
Adobe Portable Document Format
Bloque de licencias
Mostrando 1 - 1 de 1
No hay miniatura disponible
Nombre:
license.txt
Tamaño:
347 B
Formato:
Item-specific license agreed upon to submission
Descripción: