用认知诊断法让AI考生更像真人,提升教育测评校准效果
Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
- 给大模型注入多样认知能力画像,模拟真实考生答题模式
- 在536名真人数据上,认知匹配度达0.92~0.98,题目难度恢复准确率显著提升
- 适合教育测评开发、智能题库构建等需真实考生行为的场景
教育测评的标准化校准通常依赖昂贵的人类作答数据。大语言模型(LLM)可模拟考生,但其作答过于精准且同质化。本文提出零样本框架认知诊断剖面(CDP),通过自然语言生成二元属性掌握模式,并在无信息或有信息分布下采样。基于塔苏卡分数减法数据集(536名考生,15道题,5个属性),评估了八种LLM配置在无剖面、无信息CDP和有信息CDP条件下的表现。结果表明:分布重叠度提升;剖面得分与人类预期的相关性达0.92~0.98;题目难度恢复效果改善,尤其在推理增强模型中表现突出。最强案例中,Gemini 3.0 Flash(Thinking)在一参数逻辑模型下,题目难度斯皮尔曼相关系数从0.24升至0.86和0.90,均方根误差从6.31降至1.30和0.90。有信息条件对剖面一致性高的情况帮助最大。该方法使大模型模拟考生更贴近真实考生,具备实际测试开发应用价值。
原文摘要 · Abstract (English)
Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natural-language profiles and sampled under an uninformative or an informative distribution. Using the Tatsuoka fraction-subtraction dataset (536 examinees, 15 items, five attributes), we evaluated eight LLM configurations under no-profile, uninformative-CDP, and informative-CDP conditions, assessing alignment with human examinees at the ability-distribution, mastery-profile, and item-difficulty levels. CDP improved all three levels: distributional overlap rose across configurations; weighted correlations between profile-level scores and human profile expectations reached 0.92 to 0.98; and item-difficulty recovery improved in rank order and absolute alignment, most for reasoning-enabled models; in the strongest case, Gemini 3.0 Flash (Thinking), one-parameter logistic (1PL) difficulty Spearman correlations rose from 0.24 to 0.86 and 0.90 and the root-mean-square error (RMSE) fell from 6.31 to 1.30 and 0.90; the informative condition helped most where profile-level alignment was strong. CDP brings LLM-simulated examinees into closer psychometric alignment with human examinees, making them practical for operational test development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。