arXiv:2608.12652cs.CLcs.LG2026-08

提出新方法检测基准数据污染,发现现有方法存在高误报风险。

Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection

  • 通过残差流探针分析深度特征,设计零和对比协议检测污染
  • 实测显示修正项方差远超原假设,导致结果不可靠,误报率达0.72
  • 方法在真实模型上失效,提示需补充代理变量与扩大对照组

基准数据污染通常通过n-gram重叠、似然成员推断或藏宝字符串检测,但这些方法常依赖训练语料、理想统计量或发布前预判。近期一种线性探针方法可直接从内部激活中读取污染信号。本文指出自然实现方式无效,提出一种经测量验证的修正协议:对探针准确率深度分布进行零和对比,以匹配水平的伪基线为中心,对抗标签置换零假设,参考集大小为可疑集两倍。报告超出分离度而非形状,使假阳性率随分析者控制集规模变化(真零假设下0.03至0.99)。当表面解码能力随深度上升时,对比平坦深度分布的检验误报率达0.72;若其下降则完全丧失效能。真实Transformer模型测试中该协议失败,而模拟未预见此问题。因模拟中表面键是驱动变异的协变量,而真实文本中仅为代理变量;降低模拟键质量可重现此现象。引入伴随测量与扩大的零假设后,所有组均未检测到污染。

原文摘要 · Abstract (English)

Benchmark contamination is diagnosed with n-gram overlap, likelihood-based membership inference, or canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at release. A recent alternative reads it off a linear probe on internal activations. We show the natural way to do this does not work, specify one that survives measurement, then find that the correction making it work carries more variance than the null it is tested against. The protocol reports a zero-sum contrast on the depth profile of probe accuracy, recentred on a level-matched placebo baseline, tested against a label-permutation null, with the reference set twice the size of the suspect set. Each choice replaces a simpler alternative we rejected on measurement. Reporting the level of excess separability rather than its shape makes the false positive rate track the size of the analyst's own control set, 0.03 to 0.99 under a true null. Contrasting against a flat depth profile rejects a true null 0.72 of the time when surface decodability rises with depth, and loses all power when it falls. On real transformers the protocol fails a test the simulations did not pose. The recentring subtracts an estimate, and the permutation null holds it fixed. Re-estimated across split seeds on four audits of contaminated checkpoints, its standard deviation is 1.30 to 1.56 times the null's own in every arm: what is subtracted to remove a bias is more variable than what it corrects. The one nominally significant result, p = 0.0075, becomes 0.0745 once that variance is propagated, and no verdict is issued. The simulations missed this because their surface key is the covariate driving item variation; on real text it is a proxy, and degrading key quality in simulation reproduces it. We add a companion measurement and a widened null. No arm shows contamination.

模型审计数据污染探针分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。