复杂评估中,大模型评委易受隐含信息干扰,反而更不靠谱。
Curse of Knowledge: When Complex Evaluation Context Benefits yet Biases LLM Judges
- 构建ComplexEval基准,系统测试大模型在复杂任务中的评估偏差
- 所有模型均受显著偏差影响,且任务越复杂偏差越大
- 推理能力强的模型反而更易被干扰,适合评估研究者关注
随着大语言模型能力提升,其面临越来越多样且复杂的任务,可靠评估变得困难。以LLM作为评估者已成为一种可扩展的解决方案,但现有研究多聚焦于简单场景。在涉及多维度评分标准、非结构化参考答案和细微判断标准的复杂任务中,其可靠性仍缺乏研究。本文构建了ComplexEval基准,系统暴露并量化辅助信息引发的偏差。在12个基础与3个高级场景中,系统检验并验证了6种此前未被探索的偏差。关键发现包括:(1)所有被测模型均显著易受这些偏差影响,偏差程度随任务复杂度增加而上升;(2)值得注意的是,大型推理模型(LRMs)表现出反直觉的脆弱性。深入分析为提升评估信号的准确性和可验证性提供了重要洞见,推动更通用、鲁棒的评估体系发展。
原文摘要 · Abstract (English)
As large language models (LLMs) grow more capable, they face increasingly diverse and complex tasks, making reliable evaluation challenging. The paradigm of LLMs as judges has emerged as a scalable solution, yet prior work primarily focuses on simple settings. Their reliability in complex tasks--where multi-faceted rubrics, unstructured reference answers, and nuanced criteria are critical--remains understudied. In this paper, we constructed ComplexEval, a challenge benchmark designed to systematically expose and quantify Auxiliary Information Induced Biases. We systematically investigated and validated 6 previously unexplored biases across 12 basic and 3 advanced scenarios. Key findings reveal: (1) all evaluated models exhibit significant susceptibility to these biases, with bias magnitude scaling with task complexity; (2) notably, Large Reasoning Models (LRMs) show paradoxical vulnerability. Our in-depth analysis offers crucial insights for improving the accuracy and verifiability of evaluation signals, paving the way for more general and robust evaluation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。