arXiv:2601.08843cs.CLcs.LG2026-01被引 5

用大模型评分需警惕复杂评分标准下的偏差和攻击脆弱性

Rubric-Conditioned LLM Grading: Alignment, Uncertainty, and Robustness

  • 基于评分标准设计评分机制,评估大模型在不同复杂度下的评分一致性
  • 低置信度预测过滤后准确率提升,但精细评分任务效果下降
  • 对同义词替换敏感,提示注入攻击则较鲁棒,适合教育场景部署

自动短答案评分(ASAG)因学生回答的语言差异及评分标准的细致要求而面临挑战。尽管大语言模型(LLM)提供了有前景的解决方案,其在基于评分标准的评分任务中的可靠性仍需严格评估。本文系统评估了基于评分标准的LLM评分表现,重点考察三方面:在不同评分标准复杂度下,模型评分与专家判断的一致性;通过共识式弃权机制实现的置信度与准确率之间的权衡;以及在随机输入扰动和对抗攻击下的鲁棒性。使用SciEntsBank基准和Qwen 2.5-72B模型,发现二元任务中一致性较强,但评分粒度增加时性能下降。'可信度曲线'分析表明,剔除低置信度预测可提升剩余样本准确率。此外,鲁棒性实验显示模型对提示注入具有抵抗力,但对同义词替换敏感。研究揭示了评分条件化LLM裁判的能力与局限,强调不确定性估计和鲁棒性测试对可靠部署的重要性。

原文摘要 · Abstract (English)

Automated short-answer grading (ASAG) remains a challenging task due to the linguistic variability of student responses and the need for nuanced, rubric-aligned partial credit. While Large Language Models (LLMs) offer a promising solution, their reliability as automated judges in rubric-based settings requires rigorous assessment. In this paper, we systematically evaluate the performance of LLM-judges for rubric-based short-answer grading. We investigate three key aspects: the alignment of LLM grading with expert judgment across varying rubric complexities, the trade-off between uncertainty and accuracy facilitated by a consensus-based deferral mechanism, and the model's robustness under random input perturbations and adversarial attacks. Using the SciEntsBank benchmark and Qwen 2.5-72B, we find that alignment is strong for binary tasks but degrades with increased rubric granularity. Our "Trust Curve" analysis demonstrates a clear trade-off where filtering low-confidence predictions improves accuracy on the remaining subset. Additionally, robustness experiments reveal that while the model is resilient to prompt injection, it is sensitive to synonym substitutions. Our work provides critical insights into the capabilities and limitations of rubric-conditioned LLM judges, highlighting the importance of uncertainty estimation and robustness testing for reliable deployment.

自动评分大模型评估评分一致性鲁棒性测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。