arXiv:2507.08487cs.CLcs.AI2025-07

用项目反应理论优化作文连贯性评分,提升自动化评估准确率。

Enhancing Essay Cohesion Assessment: A Novel Item Response Theory Approach

  • 引入项目反应理论建模作文评分的难度与区分度
  • 在6563篇巴西高考作文上表现优于传统机器学习模型
  • 适合教育AI领域需精准评估写作质量的研究者

作文是评估写作学习成果的重要工具。文本连贯性是文本的核心特征,有助于各部分意义的建立。自动评分连贯性在教育人工智能领域仍具挑战,现有机器学习算法通常忽略语料中实例的个体特性。本文提出一种基于项目反应理论的连贯性评分预测方法,用于校准机器学习模型生成的分数。实验采用扩展版Essay-BR(6,563篇巴西高考风格作文)和巴西葡萄牙语叙事作文数据集(1,235篇来自公立学校5至9年级学生)。提取325个语言学特征,将问题视为回归任务。结果表明,所提方法在多个评价指标上优于传统机器学习模型及集成方法,为改进教育类作文自动评估提供了新思路。

原文摘要 · Abstract (English)

Essays are considered a valuable mechanism for evaluating learning outcomes in writing. Textual cohesion is an essential characteristic of a text, as it facilitates the establishment of meaning between its parts. Automatically scoring cohesion in essays presents a challenge in the field of educational artificial intelligence. The machine learning algorithms used to evaluate texts generally do not consider the individual characteristics of the instances that comprise the analysed corpus. In this meaning, item response theory can be adapted to the context of machine learning, characterising the ability, difficulty and discrimination of the models used. This work proposes and analyses the performance of a cohesion score prediction approach based on item response theory to adjust the scores generated by machine learning models. In this study, the corpus selected for the experiments consisted of the extended Essay-BR, which includes 6,563 essays in the style of the National High School Exam (ENEM), and the Brazilian Portuguese Narrative Essays, comprising 1,235 essays written by 5th to 9th grade students from public schools. We extracted 325 linguistic features and treated the problem as a machine learning regression task. The experimental results indicate that the proposed approach outperforms conventional machine learning models and ensemble methods in several evaluation metrics. This research explores a potential approach for improving the automatic evaluation of cohesion in educational essays.

作文评分项目反应理论连贯性评估教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。