模型安全比较结果可能因配置不同而反转,揭示了基准测试的潜在不稳定性。
SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks

- 提出可测量的配对分歧率,判断配置变化是否导致排序反转
- 实测发现任意基准上配置调整均可逆转模型优劣结论
- 适合关注大模型评估可靠性的研究者与工程师参考
基于基础模型基准的成对模型比较(如「A比B更安全」)常被当作定量结论,但其结果依赖于未明确定义的配置选择。本文提出一个有限包络命题,将可测量的成对分歧率与严格排序的配置对反转联系起来,并设计了一个带签名的评估协议,在多个广受引用的对齐基准上实现该机制。在所有测试的基准中,仅通过改变配置即可反转成对结论;该命题精准定位了这种严格反转的失败模式。
原文摘要 · Abstract (English)
Pairwise model comparisons drawn from foundation-model benchmarks ("A is safer than B") are read as quantitative verdicts but hinge on harness choices benchmark papers under-specify. We close one theory-benchmark loop on this primitive: a finite-envelope proposition tying a measurable pairwise-disagreement rate to whether the strict ordering admits a configuration-pair reversal, paired with a commit-stamped evaluation protocol that operationalises it on widely cited alignment benchmarks. On every benchmark we test, configuration choice alone can flip the pairwise verdict; the proposition isolates this strict-reversal failure mode.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。