arXiv:2503.04474cs.LGcs.CR2025-03ICLR被引 23

研究大模型评判者在真实场景下的可靠性,发现其易受提示和攻击影响。

Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges

  • 通过分析常见安全评判模型,揭示其对提示风格敏感
  • 微小输出风格变化可致误判率上升0.24,攻击可让100%有害内容被误判为安全
  • 适合关注大模型评估可信度的研究者和开发者

基于大语言模型(LLM)的评判者是离线基准测试、自动化红队攻击和在线防护等关键安全评估流程的基础。这一广泛依赖引发核心问题:我们能否信任这些评判者的评估结果?本文揭示两个常被忽视的关键挑战:(i) 真实场景中的评估,如提示敏感性和分布偏移可能影响性能;(ii) 针对评判模型的对抗攻击。通过对常用安全评判模型的研究发现,仅输出风格的微小变化就可能导致同一数据集上误判率上升0.24;而针对模型生成内容的对抗攻击,可使部分评判者将100%的有害生成误判为安全。这些发现暴露了现有元评估基准的缺陷以及当前LLM评判者鲁棒性的薄弱环节,表明某些评判下低攻击成功率可能带来虚假的安全感。

原文摘要 · Abstract (English)

Large Language Model (LLM) based judges form the underpinnings of key safety evaluation processes such as offline benchmarking, automated red-teaming, and online guardrailing. This widespread requirement raises the crucial question: can we trust the evaluations of these evaluators? In this paper, we highlight two critical challenges that are typically overlooked: (i) evaluations in the wild where factors like prompt sensitivity and distribution shifts can affect performance and (ii) adversarial attacks that target the judge. We highlight the importance of these through a study of commonly used safety judges, showing that small changes such as the style of the model output can lead to jumps of up to 0.24 in the false negative rate on the same dataset, whereas adversarial attacks on the model generation can fool some judges into misclassifying 100% of harmful generations as safe ones. These findings reveal gaps in commonly used meta-evaluation benchmarks and weaknesses in the robustness of current LLM judges, indicating that low attack success under certain judges could create a false sense of security.

大模型评测安全评估对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。