arXiv:2502.17785cs.CLcs.HC2025-02被引 8

用大模型自动评估阅读理解题难度,提升教育系统个性化能力

Exploring the Potential of Large Language Models for Estimating the Reading Comprehension Question Difficulty

  • 用GPT-4o和o1模型分析题目回答准确率与难度分类
  • 模型估计难度与IRT参数显著相关,但对极端题目不敏感
  • 适合用于自适应教学系统,实现大规模动态评估

阅读理解是个人成功的关键,但传统方法如语言分析和项目反应理论(IRT)需大量人工标注和大规模测试,难以扩展。大语言模型(LLMs)有望自动化难度评估,但该领域仍待深入。本研究使用SARA数据集,评估OpenAI的GPT-4o和o1在阅读理解题难度估计上的表现,既考察模型答题准确率,也评估其按IRT定义分类难度的能力。结果表明,模型估计的难度与推导出的IRT参数具显著一致性,但在极端题目特征上敏感度不足。这说明LLMs可作为可扩展的自动化难度评估工具,尤其适用于学习者与自适应教学系统(AIS)的动态交互,弥合传统心理测量方法与现代AIS之间的差距,推动更个性化、自适应的阅读评估发展。

原文摘要 · Abstract (English)

Reading comprehension is a key for individual success, yet the assessment of question difficulty remains challenging due to the extensive human annotation and large-scale testing required by traditional methods such as linguistic analysis and Item Response Theory (IRT). While these robust approaches provide valuable insights, their scalability is limited. There is potential for Large Language Models (LLMs) to automate question difficulty estimation; however, this area remains underexplored. Our study investigates the effectiveness of LLMs, specifically OpenAI's GPT-4o and o1, in estimating the difficulty of reading comprehension questions using the Study Aid and Reading Assessment (SARA) dataset. We evaluated both the accuracy of the models in answering comprehension questions and their ability to classify difficulty levels as defined by IRT. The results indicate that, while the models yield difficulty estimates that align meaningfully with derived IRT parameters, there are notable differences in their sensitivity to extreme item characteristics. These findings suggest that LLMs can serve as the scalable method for automated difficulty assessment, particularly in dynamic interactions between learners and Adaptive Instructional Systems (AIS), bridging the gap between traditional psychometric techniques and modern AIS for reading comprehension and paving the way for more adaptive and personalized educational assessments.

大模型阅读理解自适应教育难度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。