测试大模型答题是否像真人,发现需调参才更可信。
Do LLMs Give Psychometrically Plausible Responses in Educational Assessments?
- 用心理测量学框架评估18个大模型答题行为
- 大模型经温度调节后更接近人类答题分布
- 阅读题表现最好,但整体相关性仍较弱
了解测试参与者如何回答教育测评题目对题项质量评估和测验效度提升至关重要。传统方法依赖大量真人预测试,耗时费力。若大语言模型(LLMs)能表现出类人答题行为,则可作为虚拟被试加速测验开发。本文基于经典测验理论与项目反应理论,使用两个公开的多项选择题数据集(涵盖阅读、美国历史、经济学三科),评估18个指令微调的LLMs的答题表现。结果显示,更大模型往往过度自信,但通过温度缩放校准后,其答题分布更趋近人类;在阅读理解题上,模型与人类的相关性更高,但在其他科目中整体相关性较弱。因此,目前不建议在零样本场景下直接用大模型替代真人进行教育测评预测试。
原文摘要 · Abstract (English)
Knowing how test takers answer items in educational assessments is essential for test development, to evaluate item quality, and to improve test validity. However, this process usually requires extensive pilot studies with human participants. If large language models (LLMs) exhibit human-like response behavior to test items, this could open up the possibility of using them as pilot participants to accelerate test development. In this paper, we evaluate the human-likeness or psychometric plausibility of responses from 18 instruction-tuned LLMs with two publicly available datasets of multiple-choice test items across three subjects: reading, U.S. history, and economics. Our methodology builds on two theoretical frameworks from psychometrics which are commonly used in educational assessment, classical test theory and item response theory. The results show that while larger models are excessively confident, their response distributions can be more human-like when calibrated with temperature scaling. In addition, we find that LLMs tend to correlate better with humans in reading comprehension items compared to other subjects. However, the correlations are not very strong overall, indicating that LLMs should not be used for piloting educational assessments in a zero-shot setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。