检验信号检测中选择性预测的可靠性,发现多数方法实际风险远超承诺。
False Sense of Safety in Selective Signal Classification: Auditing Bound Tightness and Exchangeability for Risk Control
- 用四种校准规则测试异常声音与生成图像检测的置信边界
- 73%情况下基础方法超支预算,而严谨方法在可交换数据下零超限
- 跨类型部署时认证方法仍会失效,需分组处理但牺牲覆盖率
带有分布无关风险控制的选择性预测承诺:在置信度1-δ下,被接受输入的错误率不超过用户设定的预算α。本文针对信号域检测器(机器异常声音检测与AI生成图像取证)评估四种校准规则——未经认证的经验阈值(NAIVE)及经认证的Hoeffding、Clopper-Pearson(CP)、投注法(WSR)上界。结果显示:(i) NAIVE阈值在49%-73%的合成实验(n=200校准点)和最高68%的真实数据划分中超过预算,造成虚假安全感,因该规则本无证明;(ii) 边界紧致性至关重要:CP与WSR在可交换划分下提供显著覆盖,而Hoeffding无覆盖,且零超限;(iii) 在分组部署(未见机器类型或生成器)下,认证方法在9%-30%实验中仍超限,远高于δ,问题源于交换性假设失效而非边界本身;采用每组保守阈值可恢复有效性,但代价是严重降低覆盖率。
原文摘要 · Abstract (English)
Selective prediction with distribution-free risk control promises that, with confidence 1-delta over the calibration draw, the error rate of accepted inputs stays below a user budget alpha. We audit this promise on signal-domain detectors -- machine anomalous-sound detection (ASD) and AI-generated-image forensics -- for four calibration rules: uncertified empirical thresholding (NAIVE) and certified Hoeffding, Clopper-Pearson (CP), and betting (WSR) upper confidence bounds. We report three findings. (i) NAIVE thresholding, common in practice, exceeds its declared budget in 49-73% of synthetic trials (n=200 calibration points) and in up to 68% of real-data splits: a false sense of safety rather than a broken theorem, since the rule never had a certificate. (ii) Tightness matters: CP and WSR certify substantial coverage where Hoeffding certifies none, with zero observed budget overruns under exchangeable splits. (iii) Under grouped deployment (unseen machine types or generators), certified rules overrun in 9-30% of trials -- far above delta -- showing the failure lies in the broken exchangeability premise, not in the bounds; a conservative per-group threshold restores validity at a severe coverage cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。