arXiv:2606.05384cs.AIcs.CL2026-06中稿 · ACL

LLM判官的评估结果会因后续对话被轻易改变,影响基准测试可靠性。

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges

  • 通过后续对话挑战初始判断,发现评估结果可被逆转
  • 在MT-Bench和AlpacaEval上,80%以上判断可被动机性交互推翻
  • 适合关注评估系统鲁棒性的研究者与开发者参考

LLM作为评判者广泛用于模型评测,通常假设评判结果对固定输入是稳定的。我们发现这一假设在交互场景下不成立。研究了决策后的可操纵性:即在初始判断后,通过后续对话能否改变评估结果。在MT-Bench和AlpacaEval上的控制实验表明,重复中立重评时判断稳定,但受针对性挑战后显著可逆。反基线挑战协议显示,稳定判断可被有动机的交互推翻;平衡验证协议则分离出这种可逆性并非源于整体定向引导。这些逆转带来实际后果:降低与人类偏好的一致性,改变基准排名,产生有害评价变更,即使模型自信度很高。权威框架特别加剧不稳定性,且修正后的理由重叠度低,暗示事后合理化而非可靠纠错。我们提出评估鲁棒性分数(ERS),结合可逆性与平衡方向效应来量化交互鲁棒性。研究揭示决策后交互是LLM评判评估的一个独立失败模式,呼吁评测协议不仅衡量静态一致性,还需考察抗挑战能力。

原文摘要 · Abstract (English)

LLM-as-judge evaluation is widely used in benchmarking pipelines, where model outputs are compared and ranked using automated evaluators. These pipelines typically assume that judgments are stable properties of fixed inputs. We show that this assumption does not hold under interaction. We study post-decision manipulability: the extent to which an evaluation outcome can be altered through subsequent conversation with the judge after an initial decision has been made. Across controlled experiments on MT-Bench and AlpacaEval, we find that LLM judges are highly stable under repeated and neutral reevaluation, yet become substantially reversible under targeted post-decision challenge. An anti-baseline challenge protocol shows that stable judgments can be overturned through motivated interaction, while a counterbalanced target-validation protocol separates this reversibility from net target-directed steering. These reversals have practical consequences: they can degrade agreement with human preferences, shift benchmark rankings, and produce harmful evaluation changes despite high self-reported confidence. Authority framing is especially destabilizing, and revised judgments are often accompanied by low-overlap justifications, suggesting post hoc rationalization rather than reliable error correction. We introduce the Evaluation Robustness Score (ERS) to quantify interactional robustness by combining reversal susceptibility with counterbalanced directional effects. Our findings identify post-decision interaction as a distinct failure mode for LLM-as-judge evaluation and motivate evaluation protocols that measure not only static agreement, but robustness under challenge.

LLM评估鲁棒性可操纵性评测协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。