arXiv:2507.05129cs.CLcs.CY2025-07EMNLP被引 25

用模拟学生预测题目难度,无需真实答题数据。

SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction

  • 用偏好优化让模拟学生能力与目标一致
  • 在两个真实数据集上表现优于现有方法
  • 适合冷启动题目难度评估场景

题目难度在教育评估中至关重要,可实现对学生能力的精准测评并推动个性化学习。传统方法需真实学生答题后拟合项目反应理论(IRT)模型来估计难度,成本高且无法用于新题目的冷启动场景。本文提出SMART(基于IRT对齐的模拟学生),通过直接偏好优化(DPO)将模拟学生的能力与预设目标对齐,进而生成大量答题响应,利用大语言模型评分系统评估,并拟合IRT模型获得题目难度估计。在两个真实学生作答数据集上的实验表明,SMART因具备更优的能力对齐,在题目难度预测上显著优于其他方法。

原文摘要 · Abstract (English)

Item (question) difficulties play a crucial role in educational assessments, enabling accurate and efficient assessment of student abilities and personalization to maximize learning outcomes. Traditionally, estimating item difficulties can be costly, requiring real students to respond to items, followed by fitting an item response theory (IRT) model to get difficulty estimates. This approach cannot be applied to the cold-start setting for previously unseen items either. In this work, we present SMART (Simulated Students Aligned with IRT), a novel method for aligning simulated students with instructed ability, which can then be used in simulations to predict the difficulty of open-ended items. We achieve this alignment using direct preference optimization (DPO), where we form preference pairs based on how likely responses are under a ground-truth IRT model. We perform a simulation by generating thousands of responses, evaluating them with a large language model (LLM)-based scoring model, and fit the resulting data to an IRT model to obtain item difficulty estimates. Through extensive experiments on two real-world student response datasets, we show that SMART outperforms other item difficulty prediction methods by leveraging its improved ability alignment.

题目难度模拟学生IRTLLM评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。