arXiv:2605.25492cs.LG2026-05被引 15

模型安全比较结果可能因配置不同而反转,揭示了基准测试的潜在不稳定性。

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks

论文配图:SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks
图 1 · 摘自论文原文
  • 提出可测量的配对分歧率,判断配置变化是否导致排序反转
  • 实测发现任意基准上配置调整均可逆转模型优劣结论
  • 适合关注大模型评估可靠性的研究者与工程师参考

基于基础模型基准的成对模型比较(如「A比B更安全」)常被当作定量结论,但其结果依赖于未明确定义的配置选择。本文提出一个有限包络命题,将可测量的成对分歧率与严格排序的配置对反转联系起来,并设计了一个带签名的评估协议,在多个广受引用的对齐基准上实现该机制。在所有测试的基准中,仅通过改变配置即可反转成对结论;该命题精准定位了这种严格反转的失败模式。

原文摘要 · Abstract (English)

Pairwise model comparisons drawn from foundation-model benchmarks ("A is safer than B") are read as quantitative verdicts but hinge on harness choices benchmark papers under-specify. We close one theory-benchmark loop on this primitive: a finite-envelope proposition tying a measurable pairwise-disagreement rate to whether the strict ordering admits a configuration-pair reversal, paired with a commit-stamped evaluation protocol that operationalises it on widely cited alignment benchmarks. On every benchmark we test, configuration choice alone can flip the pairwise verdict; the proposition isolates this strict-reversal failure mode.

模型评估安全对齐基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。