Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Revisiting judging reliability in taekwondo freestyle Poomsae: implications for AI-supported evaluation.

Created on 22 Sep 2026

Authors

Min-Woo Jeon, Hong-Suk Kim, Ji-Yong Park, Hye-Soo Cho

Published in

Frontiers in psychology. Volume 17. Pages 1881296. Epub Sep 07, 2026.

Abstract

Before artificial intelligence-based judging systems can be applied to Taekwondo Poomsae, it is necessary to understand which performance components human judges can identify consistently and which remain difficult to evaluate reliably. This study examined reliability in freestyle Poomsae judging to identify evaluation components with stable and unstable human scoring.
Ten internationally certified Taekwondo Poomsae referees evaluated ten official competition videos twice, separated by a one-week washout period. Systematic session effects were evaluated with separate linear mixed-effects models estimated by restricted maximum likelihood using the SPSS MIXED procedure. Session was specified as a fixed effect, judge and video as crossed random intercepts, and the two observations within each judge-video pair as repeated measurements with an unstructured residual covariance matrix. Fixed-effect inference used Satterthwaite denominator degrees of freedom. Test-retest stability was evaluated for identical judge-video pairs, with primary emphasis on the absolute-agreement ICC because systematic session effects were detected. Inter-rater agreement was evaluated separately by session using two-way random effects, single rating absolute-agreement and consistency ICCs with 95% confidence intervals and Kendall's W. Item-specific analyses were exploratory, and no multiplicity adjustment was applied.
The total score increased by 0.304 points, 95% CI [0.187, 0.421], t (99) = 5.177, p < 0.001. At the nominal 0.05 level, exploratory item-specific increases were observed for jumping side kick (β = 0.057, p < 0.001), acrobatic kicking technique (β = 0.021, p = 0.003), and expression of energy (β = 0.047, p = 0.003); the increase for basic movements and practicability did not reach the 0.05 threshold (p = 0.053). Absolute-agreement test-retest ICCs ranged from 0.039 to 0.848 and were excellent for harmony and the total score. All single-rating inter-rater ICC(A,1) point estimates were below 0.40. For the total score, ICC(A,1) was -0.007, 95% CI [-0.018, 0.036], in the first session and 0.060, 95% CI [0.010, 0.224], in the second session.
Systematic session effects, absolute test-retest agreement, relative consistency over time, and inter-rater agreement represent distinct aspects of judging reliability. Evaluation components with poor single-judge agreement require clearer operational definitions and further validation before computational or AI-assisted scoring is considered.

PMID:
42769074
Bibliographic data and abstract were imported from PubMed on 22 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 5
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement