An Exploratory Evaluation of GPT-4's Consistency as an English Essay Rater: A Many-Facet Rasch Model Analysis of AI versus Human Rating Patterns
Austin Pack
Steven Carter
Brigham Young University-Hawaii, USA
Alex Barrett
Florida State University, USA
Juan Escalante
Brigham Young University-Hawaii, USA
Mark Wolfersberger
Brigham Young University, USA
Abstract
This study examined the defensibility of using GPT-4 for automated essay scoring, using a ManyFacet Rasch Model analysis. Forty English for academic purposes student essays were rated by GPT- 4 and four trained educators to assess nuances in rubric application, severity, leniency, and bias. Findings suggest that while GPT-4 tended to avoid the use of extreme scores, exhibiting a moderate central tendency rating, it does show a high level of consistency in its scoring behavior. This study contributes to understanding the extensions and limitations of using Generative AI tools in scoring essays, and provides insights into the use of AI tools in assessing writing.
Keywords
Generative AI, ChatGPT, artificial intelligence, automated essay scoring, assessment, education
