用心理测量学方法分析大模型评分能力,发现难答对错集中于中间标签。
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory

- 基于项目反应理论建模评分准确率与答题难度、模型能力的关系
- 17个模型在难答上表现差异大,准确率下降速度不一
- 难答特征:语义偏离参考答案、矛盾信号强、嵌入空间孤立
大语言模型(LLM)在自动短答案评分中的评估常依赖宏平均F1和Cohen's kappa等综合指标,但这些指标难以揭示评分性能随答题难度变化的细节。本文提出一种基于项目反应理论(IRT)的评估框架,将评分正确性建模为隐含评分者能力与答题难度的函数。该框架可实现响应级别的分析,揭示仅从整体得分无法观察到的鲁棒性差异。我们在SciEntsBank和Beetle基准上对17个开源权重的LLM应用此框架。结果显示,即使整体性能相似,各模型在答题难度上升时准确率下降速度也显著不同。此外,错误模式显示,对困难回答的误判高度集中在\texttt{partially_correct_incomplete}标签,表明模型在模糊情境下存在中间标签坍塌倾向。为进一步刻画难点,我们分析了估计难度的语义与语言学相关因素。在两个数据集上,高难度均与参考答案语义对齐度低、矛盾信号强、嵌入空间中语义孤立性高有关。总体而言,项目反应理论为超越综合指标的LLM评分评估提供了有效框架。
原文摘要 · Abstract (English)
Automated short answer grading (ASAG) with large language models (LLMs) is commonly evaluated with aggregate metrics such as macro-F1 and Cohen's kappa. However, these metrics provide limited insight into how grading performance varies across student responses of differing grading difficulty. We introduce an evaluation framework for LLM-based ASAG based on item response theory (IRT), which models grading correctness as a function of latent grader ability and response grading difficulty. This formulation enables response-level analysis of where LLM graders succeed or fail and reveals robustness differences that are not visible from aggregate scores alone. We apply the framework to 17 open-weight LLMs on the SciEntsBank and Beetle benchmarks. The results show that even models with similar overall performance differ substantially in how sharply their grading accuracy declines as response difficulty increases. In addition, confusion patterns show that errors on difficult responses concentrate disproportionately on the \texttt{partially\_correct\_incomplete} label, indicating a tendency toward intermediate-label collapse under ambiguity. To characterize difficult responses, we further analyze semantic and linguistic correlates of estimated difficulty. Across both datasets, higher difficulty is associated with weaker semantic alignment to the reference answer, stronger contradiction signals, and greater semantic isolation in embedding space. Overall, these results show that item response theory offers a useful framework for evaluating LLM-based ASAG beyond aggregate performance measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。