arXiv:2503.09347cs.CLcs.AI2025-03ACL被引 27

大模型当安全评估员易被文字套路干扰,结果不可靠。

Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts

  • 用多个大模型做安全评估,看它们是否一致
  • 道歉或啰嗦的表达能让判断偏移高达98%
  • 合议制也难根除干扰,需更抗干扰方法

大型语言模型(LLMs)正被广泛用作自动化安全评估工具,但其可靠性仍存疑。本研究在关键安全领域评估了11种LLM评判模型,重点考察三方面:重复判断的一致性、与人类判断的对齐度,以及对输入中道歉或冗长表述等伪饰文本的敏感性。结果显示,评测模型中的偏差会显著扭曲内容安全性判断,导致评估失效。特别地,仅道歉类语言即可使模型偏好产生高达98%的偏差。出人意料的是,模型越大并不一定更稳健,小模型反而在某些情况下更具抗干扰能力。为提升鲁棒性,研究尝试使用多模型合议机制,虽能增强一致性与人类对齐度,但对伪饰文本的敏感性依然存在,即使最佳合议配置亦无法完全消除。这表明亟需开发多样化、抗伪饰的安全评估方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly employed as automated evaluators to assess the safety of generated content, yet their reliability in this role remains uncertain. This study evaluates a diverse set of 11 LLM judge models across critical safety domains, examining three key aspects: self-consistency in repeated judging tasks, alignment with human judgments, and susceptibility to input artifacts such as apologetic or verbose phrasing. Our findings reveal that biases in LLM judges can significantly distort the final verdict on which content source is safer, undermining the validity of comparative evaluations. Notably, apologetic language artifacts alone can skew evaluator preferences by up to 98\%. Contrary to expectations, larger models do not consistently exhibit greater robustness, while smaller models sometimes show higher resistance to specific artifacts. To mitigate LLM evaluator robustness issues, we investigate jury-based evaluations aggregating decisions from multiple models. Although this approach both improves robustness and enhances alignment to human judgements, artifact sensitivity persists even with the best jury configurations. These results highlight the urgent need for diversified, artifact-resistant methodologies to ensure reliable safety assessments.

大模型评估安全评测伪饰攻击合议机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。