用微调大模型模拟学生作答,生成题目特征曲线。
Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models
- 微调大模型根据能力水平生成答题行为,隐式建模题目难度与区分度。
- 在六年级英语题集上,预测正确率曲线与真实数据接近,区分度建模更优。
- 无需真实测试数据,适合快速构建评估题库,适合教育AI研发者。
传统评估题参数(如难度、区分度)的确定依赖昂贵的真实考试数据收集以进行项目反应理论(IRT)校准。本文提出一种新方法:通过微调大语言模型(LLMs),模拟不同潜质学生对多项选择题的作答行为,隐式建模心理测量属性。基于Qwen-3密集模型系列与低秩适配(LoRA),训练模型在离散能力描述条件下生成回答,重构正确作答概率随学生能力变化的函数,从而生成合成的项目特征曲线(ICC),用于估计IRT参数。在六年级英语语言艺术(ELA)题集和BEA 2024共享任务数据集上的评估表明,该方法性能可媲美或优于基线方法,尤其在建模题目区分度方面表现突出。
原文摘要 · Abstract (English)
Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibration. This study introduces a novel approach that implicitly models these psychometric properties by fine-tuning Large Language Models (LLMs) to simulate student responses across a spectrum of latent abilities. Leveraging the Qwen-3 dense model series and Low-Rank Adaptation (LoRA), we train models to generate responses to multiple choice questions conditioned on discrete ability descriptors. We reconstruct the probability of a correct response as a function of student ability, effectively generating synthetic Item Characteristic Curves (ICCs) to estimate IRT parameters. Evaluation on a dataset of Grade 6 English Language Arts (ELA) items and the BEA 2024 Shared Task dataset demonstrates that this method competes with or outperforms baseline approaches. This simulation-based technique seems particularly effective at modeling item discrimination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。