仅凭模型输出就能检测数据泄露,且无需先验知识。
A prior-free blind detection of information leakage from model predictions

- 基于预测值与结果的分布关系设计无先验检测方法
- 发现近确定性子群体是泄露的唯一可检测特征,阈值差约0.007
- 适合审计模型输出的科研人员或合规团队使用
数据泄露——模型被包含基线不可用信息污染——是机器学习驱动科学中复现失败的主要原因,但现有检测工具需训练代码、外部数据或领域知识。而审计者最常持有的只有模型输出。本文研究仅从预测结果和真实结果中能判断什么。提出一种决策理论框架,将泄露诊断定义为预测风险/结果分布的泛函,由与正确评分规则和决策曲线分析相关的阈值加权参数化。证明一个严格不可能性:一个校准后与诚实模型相同、判别力也匹配的泄漏模型,无法被任何预测函数区分,因此广义泄露只能在外部提供判别力上限时才能检测。随后证明泄露无法隐藏的特征:近确定性子群体会形成持续单位纯度头部,任何合法非确定性预测器都无法生成此现象,从而实现无先验测试。结果将泄露分为三类——校准错误、宽范围校准、确定性——每类对应一个检测器和失效模式。在英国生物银行数据上验证,使用时间窗共病泄露(已知分级严重性),测得该终点检测下限Δ ext{c}^* ≈ 0.007,低于此值的残留泄露无法从输出中检测,且小到不影响结论。数值下限依赖队列与终点;结构启示具普遍性:当残留泄露与更优预测器不可区分时,输出仅检测即失效。该测试在普通硬件上对单个预测向量判决时间不足一秒。
原文摘要 · Abstract (English)
Data leakage -- contamination of a model with information unavailable at baseline -- is the dominant reproducibility failure in machine-learning-based science, yet detection tools require training code, external data, or domain expertise. None operates on the artifact an auditor most often holds: the model's output. We ask what can be decided about leakage from predictions and outcomes alone. We give a decision-theoretic framework in which leakage diagnostics are functionals of the predicted-risk/outcome law, parameterized by a threshold-weighting linked to proper scoring rules and decision-curve analysis. We prove a sharp impossibility: a recalibrated leak matching an honest model's calibration and discrimination is indistinguishable from honest performance by \emph{any} function of the predictions, so the broad class is detectable only against an externally supplied ceiling on achievable discrimination. We then prove what leakage cannot hide: a near-deterministic subgroup -- the signature of a near-label leak -- produces a sustained unit-purity head that no legitimate predictor of a non-deterministic outcome can manufacture, yielding a prior-free test. These results organize leakage into a trichotomy -- miscalibrated, broad-calibrated, and deterministic -- each with a matched detector and failure mode. We validate on UK Biobank using time-windowed comorbidity leakage with known, graded severity, measuring a detection floor of $Δ\cstar \approx 0.007$ on this endpoint, below which residual leakage is undetectable from output and too small to alter conclusions. The numerical floor is cohort- and endpoint-specific; the structural lesson is general: output-only detection fails where residual leakage is indistinguishable from an honestly stronger predictor. The test returns a verdict on a prediction vector in under a second on commodity hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。