9. Comparing Human Raters’ and ChatGPT-4’s Scores in Measuring Content in CLIL Presentation
Rie Koizumi,
Saki Suemori,
Yusuke Kubo
This chapter
is part of: Carol A. Chapelle et al. 2025. Researching Generative AI in Applied Linguistics
Description
Considering the improvement of AI in accessibility and response quality, this study examines the degree to which ChatGPT-4 generates scores comparable to human raters on content aspects in Content and Language Integrated Learning (CLIL) presentations, where both content and language instruction are prioritized, necessitating effective content assessment. We obtained 52 presentation scripts from 37 university students who had taken a CLIL course. Using a detailed prompt that includes an analytic rubric, its explanation, example scripts, and benchmarking scores, we asked ChatGPT-4 to score each presentation script. This procedure was repeated three times, and the scores were compared with human-rated scores primarily using many-facet Rasch measurement. The results showed that seven human raters and ChatGPT-4 produced consistent scores, with ChatGPT-4 tending to produce slightly lenient ratings and fewer biases. We discuss the implications of using ChatGPT-4 as a supplementary rater.
Read
more
Publication Details
Published:
August 27, 2025
Publisher: Iowa State University Digital Press
Pages: 34
DOI: 10.31274/isudp.2025.211.09
License Information:
© 2025 Koizumi, Suemori, and Kubo. Published under a CC BY license.
Citation
Koizumi, R., Suemori, S., & Kubo, Y. (2025). Comparing human raters’ and ChatGPT-4’s scores in measuring content in CLIL presentation. In C. A. Chapelle, G. H. Beckett, & B. E. Gray (Eds.), Researching generative AI in applied linguistics (pp. 166–196). Iowa State University Digital Press. https://doi.org/10.31274/isudp.2025.211.09