大模型评分在中等质量答案上表现差,需更多针对性训练来公平评估学生进展。
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation

- 用少量示例微调大模型,但对部分正确的中等答案评分不稳定。
- 人类评分最稳定,大模型在中等质量答案上准确率下降超30%。
- 针对性数据越多,评分越准,适合教育评估与模型优化研究者。
自动短答案评分正从微调模型转向少样本设置下的大语言模型(LLM),利用其广泛世界知识和部署便捷性,但任务特定数据有限可能导致复杂评分任务的对齐不足。尤其对需要精细理解的部分正确答案,影响尚不明确。我们研究不同模型的任务适应程度与其质量条件评分一致性之间的关系。在两个开放题生物题目上,对比了三种大模型(GPT-5.2、GPT-4o、Claude Opus 4.5)在少样本模式下的表现,一个微调的BERT编码器,以及一位生物教育专家,使用数百名学生的回答及专家提供的真实分数。结果显示,人类间评分一致最高且在全质量范围内稳定。所有AI模型在完全正确或完全错误的回答上表现良好,但在中等质量回答上出现显著退化。这种中段退化与任务特定适配度相关:少样本大模型在示例少时退化最严重,随着任务数据增加而缓解,微调编码器表现最佳。该现象可能造成对理解能力发展中的学生不公平评价。研究强调质量条件公平性的重要性,尤其关注中等质量回答。
原文摘要 · Abstract (English)
Automated short answer scoring (ASAS) is shifting from discriminative, fine-tuned models to large language models (LLMs) used in few-shot settings. This paradigm leverages LLMs broad world knowledge and ease of deployment, but limited task-specific data may reduce alignment on complex scoring tasks. In particular, its impact on scoring partially correct responses that require nuanced interpretation remains underexplored. We investigate the relationship between the degree of task-specific adaptation of different models and quality-conditioned scoring agreement. We compare three LLMs (GPT-5.2, GPT-4o, Claude Opus 4.5) in few-shot mode, a fine-tuned BERT-based encoder, and a human expert on two open-ended biology items, using several hundred student responses and ground truth scores provided by a biology education expert. The results show that human-human agreement is highest and stable across the full quality spectrum. All AI models perform well on fully correct and fully incorrect responses, but exhibit substantial degradation on mid-range responses. This mid-range degradation is conditioned on task-specific adaptation: It is most severe in few-shot LLMs with few examples and decreases as task-specific data increases, with fine-tuned encoder models performing best. This mid-range degradation may lead to inequitable evaluation of responses produced by students with developing understanding. Our findings highlight the importance of quality-conditioned fairness, with particular attention to mid-range responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。