arXiv:2603.06594cs.CLcs.AI2026-03被引 17

LLM判官评估安全性的可靠性存疑,真实攻击成功率被高估。

A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness

  • 用6642条人工标注验证,发现判官性能常退化到随机水平
  • 不同模型风格、攻击扭曲、语义模糊导致判官判断失准
  • 提出新基准与压力测试集,提升评估可信度

自动化LLM作为判官的框架已成为自然语言处理中可扩展评估的默认标准。例如,在安全评估中,这些判官被用来判断有害性,以衡量模型对对抗攻击的鲁棒性。然而,我们发现现有验证协议未能考虑红队测试中固有的显著分布偏移:不同受害者模型表现出不同的生成风格,攻击会扭曲输出模式,且语义模糊性在越狱场景中差异极大。通过使用6642条人工验证标签进行综合审计,我们揭示这些偏移的不可预测交互常导致判官性能退化至接近随机水平。这与先前研究中报告的高人类一致性形成鲜明对比。关键的是,我们发现许多攻击通过利用判官缺陷来虚增成功率达到效果,而非真正生成有害内容。为此,我们提出ReliableBench,一个行为更稳定可判的基准,以及JudgeStressTest,一个用于暴露判官失败的数据集。数据已公开于:https://github.com/SchwinnL/LLMJudgeReliability。

原文摘要 · Abstract (English)

Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, these judges are relied upon to evaluate harmfulness in order to benchmark the robustness of safety against adversarial attacks. However, we show that existing validation protocols fail to account for substantial distribution shifts inherent to red-teaming: diverse victim models exhibit distinct generation styles, attacks distort output patterns, and semantic ambiguity varies significantly across jailbreak scenarios. Through a comprehensive audit using 6642 human-verified labels, we reveal that the unpredictable interaction of these shifts often causes judge performance to degrade to near random chance. This stands in stark contrast to the high human agreement reported in prior work. Crucially, we find that many attacks inflate their success rates by exploiting judge insufficiencies rather than eliciting genuinely harmful content. To enable more reliable evaluation, we propose ReliableBench, a benchmark of behaviors that remain more consistently judgeable, and JudgeStressTest, a dataset designed to expose judge failures. Data available at: https://github.com/SchwinnL/LLMJudgeReliability.

LLM安全判官评估对抗鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。