Authors
Vinicius C Destefani, Matheus Feliciano C Ferreira, Afranio C Destefani
Published in
Cureus. Volume 18. Issue 8. Pages e115344. Epub Aug 28, 2026.
Abstract
Generative artificial intelligence (AI) models can create fluent medical assessment materials, including multiple-choice questions, distractors, clinical vignettes, answer explanations, and blueprint tags. However, linguistic plausibility does not establish measurement validity. Using AI-generated text as ready-made assessment material may introduce risks related to clinical accuracy, item-writing quality, blueprint alignment, distractor functioning, fairness, and score interpretation. This technical report presents a human-governed validation workflow that treats AI-generated medical assessment artifacts as candidates requiring staged review rather than immediate deployment. The workflow is organized into seven gates: artifact taxonomy, AI provenance and human accountability, clinical and content review, item-writing and language review, blueprint alignment, psychometric-readiness decision, and final human approval. The psychometric-readiness gate explicitly asks whether item analysis, distractor analysis, differential item functioning (DIF) review, or other empirical evaluation is required before an artifact is used in a scored assessment. By separating expert clinical review from psychometric evidence, the workflow preserves human accountability while acknowledging that expert approval alone cannot support all score-based interpretations. We recommend introducing AI-generated assessment materials into medical education pipelines as preliminary candidates subject to documented governance, human oversight, and empirical evaluation when used for scored or consequential assessment.
PMID:
42802837
Bibliographic data and abstract were imported from PubMed on 28 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 12
- Comments 0