arXiv:2605.06161cs.AIcs.SE2026-05被引 6

现有大模型评判系统可信度存疑,新方法可检测其判断是否受提示词扰动影响。

Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges

论文配图:Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
图 1 · 摘自论文原文
  • 提出政策不变性测试,检验评判结果是否受提示改写影响
  • 发现9.1%判决因语义不变的提示重写而改变,43%错误发生在明确案例中
  • 提供可复现的评估协议,让安全评测摆脱对评判模型盲目信任

大模型作为评判者已成为智能体安全性的标准评估工具,但现有基准将评判结果视为真实标签,未检验其是否依赖于代理行为,还是仅受评价策略表述方式影响。我们提出可信的安全评判者必须满足基本性质——政策不变性,并将其操作化为三项可测试原则:在认证等价重写下评分标准语义不变、在故意从严到从宽调整阈值时评分不变,以及对模糊情况保持敏感而稳定。通过在ASSEBench和R-Judge数据集上对四类代理评判者实施该压力测试协议,发现当前评判系统对有意义的规范变化与无意义的结构重写反应强度相当,无法区分二者。语义保持的提示重写导致高达9.1%的判决偏离基线波动,其中18%-43%的判决错误出现在本应明确的案例中,说明现有安全评分混淆了代理行为与提示方式。除诊断外,我们提出政策不变性得分与裁判卡报告协议,揭示了评判可靠性存在数量级差异,这在仅看准确率的排行榜中完全不可见。代码与协议已公开,使未来安全基准能自主审计其评估者而非默认信任。

原文摘要 · Abstract (English)

LLM-as-a-Judge pipelines have become the de facto evaluator for agent safety, yet existing benchmarks treat their verdicts as ground-truth proxies without checking whether the verdicts depend on the agent's behavior or merely on how the evaluation policy happens to be worded. We argue that any trustworthy safety judge must satisfy a basic property we call policy invariance, and we operationalize it as three testable principles: rubric-semantics invariance under certified-equivalent rewrites, rubric-threshold invariance under intentional strict-to-lenient shifts, and ambiguity-aware calibration so that verdict instability concentrates on genuinely ambiguous cases. Instantiating these principles as a stress-test protocol with four agent-class judges on trajectories drawn from ASSEBench and R-Judge, we surface a previously unmeasured failure mode: today's judges respond to meaningful normative shifts and to meaningless structural rewrites with comparable strength, and cannot tell the two apart. Content-preserving policy rewrites flip up to 9.1% of verdicts above baseline jitter, and 18-43% of all observed flips occur on unambiguous cases under such rewrites, so existing safety scores conflate what the agent did with how the evaluator was prompted. Beyond the diagnosis, we contribute the Policy Invariance Score and the Judge Card reporting protocol, which expose an order-of-magnitude spread in judge reliability that is invisible to accuracy-only leaderboards. We release the protocol and code so that future agent-safety benchmarks can audit their own evaluators rather than trust them by default.

大模型评估安全评测提示鲁棒性可信评判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。