Authors
Emily Rush, Md Nazmul Karim, George S Yacu, Jessica N Byram, Colleen N Garnett, Nicole DeVaul, Laura Smith, Margaret Checchi, Daniel Martin, Leslie A Hoffman, Kirsten M Brown, Daniel J Mumbower, Robert M Becker, Victoria A Roach, Alison F Doubleday, Danielle N Edwards, Rebecca S Lufler, Alexandra Wactor, Sophia Boxerman, Suzanne Smith, Hannah L Herriott, Megan E Kruskie, Kyle A Robertson, Elizabeth R Agosto, Christopher Facer, Abdel Metwally, Melissa Barbosa, Dahlia Chavez, Ali Akram, Truman Steele, Seth Adler, Joshua Samaniego, Sara Aqel, Chloie Flores, Yi Gao, Emily Nguyen, Melissa Petito, Adam B Wilson
Published in
Medical teacher. Pages 1-13. Aug 26, 2026. Epub Aug 26, 2026.
Abstract
Large language models (LLMs) are increasingly proposed as deductive coders in qualitative research, but their measurement properties remain underexplored. This study applies generalizability theory to evaluate whether hybrid human-LLM workflow configurations can achieve reliable mode-outcomes as an alternative consensus-generating mechanism for deductive coding tasks in medical education research.
Three commercial LLMs (GPT-5.2, Claude Opus 4.5, Gemini 3-Flash Preview) coded 741 excerpts from a published audit of AI-related policy documents at 146 U.S. medical schools against a 24-subtheme deductive framework. Mixed-effects logistic regression assessed variability in agreement with human consensus across coder type (human versus LLM), excerpt characteristics (complexity and length), and coding conditions (sequential independent versus batched processing). A simulation-based D-study forecasted agreement levels for various hybrid human-LLM configurations.
Sequential independent LLM coding showed agreement comparable to human coding (β = -0.14, p = 0.326). Batched processing showed significantly lower human-LLM agreement (β = -0.41, p = 0.007) and substantial batch-related variance (variance = 30.16). The three LLMs produced similar agreement (joint Wald χ2 = 0.15, p = 0.930), and LLM-LLM agreement (κ = 0.750-0.755) substantially exceeded human-human agreement (κ = 0.422) and human-versus-consensus agreement (κ = 0.420-0.533). LLM disagreements were more systematically patterned across the codebook (Cramer's V = 0.554-0.585) than human disagreements (Cramer's V = 0.196). Simulations forecasted agreement ranging from 0.520 to 0.529 when two or more LLMs were paired with one to two human coders. These forecasted agreement levels for hybrid human-LLM workflows fell within the range of human-versus-consensus agreement.
D-study simulations support the use of hybrid human-LLM workflows to reach coding consensus for deductive reasoning tasks through mode responses, an alternative consensus-generating mechanism to traditional adjudication discussions. Future work should examine whether these patterns extend across additional deductive coding contexts and model families.
PMID:
42647413
Bibliographic data and abstract were imported from PubMed on 27 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 30
- Comments 0