提出可检测基准污染的理论边界与校准审计方法,解决误判难题。
When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

- 基于行为差异建模污染程度,用统计矩分析可检测性条件。
- 实验证明冻结校准预估功率曲线,但小样本下预算严重失准。
- 提出两阶段审计规划,能自适应保守判断,避免误判。
行为污染检测器返回“无证据”可能因基准干净,也可能因审计能力不足。我们针对训练中未知比例 alpha 的样本曾被见过的基准进行形式化分析。在匹配的干净与已见对照下,行为通道为稀疏混合分布 Q_alpha = (1 - alpha) P_0 + alpha P_1,精确的二阶矩分析表明可检测性由 alpha * rho * sqrt(m) 决定,其中 rho^2 = chi^2(P_1 || P_0) 衡量行为可区分性。任意标量检测器的有效性 ef = |E_1 f - E_0 f| / sqrt(Var_0(f)) <= rho 可在审计前从对照中估计。独立样本分割证书可无假设地下界 alpha,无需定向假设。实证发现具有双重性:冻结校准有效性可预测保留测试的功率曲线,六种精确置换通道 R² 达 0.83–0.98;但仅靠有效性推导的高斯预算在小样本下严重失准,在 9/9 通过门控的通道中均失效,尽管有效性本身仍可传递。失败源于反演过程而非校准本身。预先声明的两阶段规划器通过模拟部署测试修复预算,保持统一保守性,并在探测信号不传递时主动放弃。证书在审计规模下有效但空洞;五种子对注入实验复现了机制排序:原文 > 重述 > 表面,说明仅答案信号实际来自基线漂移。报告审计契约及其失败:非拒绝结果只有在同时知晓有效性、预算与有效性门控的前提下才可解释。
原文摘要 · Abstract (English)
Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark in which an unknown fraction alpha of items was seen during training. With matched clean and seen controls, the behavioral channel is the sparse mixture Q_alpha = (1 - alpha) P_0 + alpha P_1, and an exact second-moment argument shows that detectability is governed by alpha * rho * sqrt(m), where rho^2 = chi^2(P_1 || P_0) measures behavioral separability. Any scalar detector reduces to its efficacy, ef = |E_1 f - E_0 f| / sqrt(Var_0(f)) <= rho, which can be estimated from controls before the audit is run. A separate sample-split certificate lower-bounds alpha distribution-free, without requiring an orientation assumption. Our empirical finding is two-sided. Frozen calibration efficacy predicts held-out power curves, with R^2 = 0.83-0.98 across six exact-permutation channels, but the efficacy-only Gaussian budget is miscalibrated at the small sample sizes it prescribes, failing in 9/9 gate-passing channels even though efficacy itself transports. The failure is in the inversion, not the calibration. A predeclared two-stage planner that simulates the deployed test repairs the budgets, is uniformly conservative, and abstains when its probe does not transport. The certificate is valid but vacuous at audit scale, and a five-seed paired injection study recovers the mechanism ordering verbatim > paraphrase > surface, in which the apparent answer-only signal is explained by baseline drift. We report the audit contract and its failures together: a non-rejection is interpretable only alongside the efficacy, budget, and validity gates that produced it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。