Authors
Francisco Cardozo, Eric C Brown, Juliana Mejía-Trujillo, Raymond Balise, Catalina Cañizares, Augusto Pérez-Gómez, Sara M St George, Vilma Gabbay
Published in
JMIR AI. Volume 5. Pages e95964. Aug 25, 2026. Epub Aug 25, 2026.
Abstract
Motivational interviewing (MI) is widely used in preventive interventions, yet coding MI techniques and monitoring intervention adherence remain resource-intensive due to the reliance on manual transcription and expert review. Large language models (LLMs) offer a promising approach to automate these tasks, but their agreement with human coders in the context of prevention interventions has not been established.
This study evaluated the agreement between an AI-based coder (OpenAI's GPT 4.1) and trained human coders on two tasks: (1) identification of MI techniques (eg, open questions, affirmations, giving information) at the facilitator-message level and (2) completing a 21-item checklist of implementation adherence for a brief MI-based preventive intervention for adolescent substance use.
Two certified MI facilitators independently coded 72 facilitator messages from 2 standardized Spanish-language Brief Intervention Based on Motivational Interviewing program (Intervención Breve Basada en Entrevista Motivacional [IBEM]) practice sessions with chatbot-simulated adolescent responses. The facilitators classified MI techniques using the OARS (open questions, affirmations, reflections, and summaries) framework and completed a 21-item implementation-adherence checklist. An AI-based coder (OpenAI's GPT-4.1, accessed through the API) classified the same facilitator messages and checklist items using a structured prompt derived from the MI coding manual. MI techniques were compared at the facilitator-message level and implementation adherence at the session level. Intercoder agreement in use of MI techniques and implementation adherence was assessed using Cohen κ, Fleiss κ, and Cochran Q tests.
For use of MI techniques, the AI coder demonstrated moderate-to-substantial agreement with human coders across most techniques, including open questions (κ=0.66-0.69), affirmations (κ=0.66-0.77), and giving information (κ=0.91). No statistically significant differences in percentages of MI technique use were observed among the 3 coders, although agreement was the weakest for higher-inference categories such as complex reflections (κ=0.00). For MI implementation adherence, overall agreement was moderate (Fleiss κ=0.487), and pairwise agreement between the AI coder and 1 human coder was substantial (κ=0.67), exceeding the agreement observed between the 2 human coders (κ=0.53).
These findings provide support for the feasibility of using LLMs to recognize MI techniques and assess implementation adherence. The results support a human-AI collaborative model in which the AI coder "precodes" facilitator messages and flags sessions for expert review, while human coders retain responsibility for higher-inference judgments and shift their effort from routine coding toward contextual review and coaching feedback. Because the analyses are based on only 2 sessions, these results should be interpreted as early-stage, proof-of-concept evidence rather than a basis for large-scale deployment. Future research should compare different LLMs and evaluate whether AI-assisted coding improves the scalability of routine implementation monitoring.
PMID:
42640265
Bibliographic data and abstract were imported from PubMed on 25 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 18
- Comments 0