arXiv:2602.00034cs.CYcs.AI2026-02被引 1

用大模型生成答题特征,无需真实测试就能估算题目难度。

Synthetic Student Responses: LLM-Extracted Features for IRT Difficulty Parameter Estimation

  • 结合语言特征与大模型提取的解题步骤、认知复杂度等新特征。
  • 在未见过的题目上预测难度,相关性达0.78,接近真实值。
  • 适合教育评估者快速设计题目,降低试测成本。

教育评估高度依赖题目难度,传统方法需耗费大量资源进行学生预测试。本文探索是否可在不进行真实学生测试的情况下,通过建模答题过程来准确估计项目反应理论(IRT)的难度参数,并分析不同特征类型对预测精度的贡献。方法结合传统语言学特征与利用大语言模型(LLMs)提取的教学洞察,包括解题步骤数量、认知复杂度及潜在误解点。采用两阶段流程:先训练神经网络预测学生对题目的作答行为,再从模拟作答模式中推导难度参数。基于超过25万份数学题目学生作答数据集,模型在完全未见题目上的预测难度与实际难度之间的皮尔逊相关系数约为0.78。

原文摘要 · Abstract (English)

Educational assessment relies heavily on knowing question difficulty, traditionally determined through resource-intensive pre-testing with students. This creates significant barriers for both classroom teachers and assessment developers. We investigate whether Item Response Theory (IRT) difficulty parameters can be accurately estimated without student testing by modeling the response process and explore the relative contribution of different feature types to prediction accuracy. Our approach combines traditional linguistic features with pedagogical insights extracted using Large Language Models (LLMs), including solution step count, cognitive complexity, and potential misconceptions. We implement a two-stage process: first training a neural network to predict how students would respond to questions, then deriving difficulty parameters from these simulated response patterns. Using a dataset of over 250,000 student responses to mathematics questions, our model achieves a Pearson correlation of approximately 0.78 between predicted and actual difficulty parameters on completely unseen questions.

教育评估大模型题难易度IRT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。