攻击可大幅提高错误率,暴露现有异常检测方法的脆弱性。
On the Adversarial Robustness of Learning-based Conformal Novelty Detection
- 设计基于查询的对抗攻击,量化错误率恶化上限
- 实验证明攻击使错误率显著上升但检测能力仍强
- 揭示当前有保证的异常检测方法的根本缺陷
本文研究基于学习的合效异常检测方法在对抗攻击下的鲁棒性。重点关注两种具备有限样本错误发现率(FDR)控制能力的先进框架:基于正负样本分类器的AdaDetect(Marandon et al., 2024)和基于一类分类器的方法(Bates et al., 2023)。尽管它们在正常条件下提供严格的统计保证,但在对抗扰动下的表现尚未充分探索。我们首先在AdaDetect框架下提出一个理想化敌手攻击设定,量化最坏情况下的FDR退化,并推导出攻击的统计代价上界。该设定直接启发了一个仅需查询输出标签的实用有效攻击方案。结合两种主流互补的黑盒对抗算法,在合成与真实数据集上系统评估了两种框架的脆弱性。结果表明,对抗扰动可在保持高检测能力的同时显著提升FDR,暴露出当前误差受控异常检测方法的根本局限,推动更鲁棒替代方案的发展。
原文摘要 · Abstract (English)
This paper studies the adversarial robustness of conformal novelty detection. In particular, we focus on two powerful learning-based frameworks that come with finite-sample false discovery rate (FDR) control: one is AdaDetect (by Marandon et al., 2024) that is based on the positive-unlabeled classifier, and the other is a one-class classifier-based approach (by Bates et al., 2023). While they provide rigorous statistical guarantees under benign conditions, their behavior under adversarial perturbations remains underexplored. We first formulate an oracle attack setup, under the AdaDetect formulation, that quantifies the worst-case degradation of FDR, deriving an upper bound that characterizes the statistical cost of attacks. This idealized formulation directly motivates a practical and effective attack scheme that only requires query access to the output labels of both frameworks. Coupling these formulations with two popular and complementary black-box adversarial algorithms, we systematically evaluate the vulnerability of both frameworks on synthetic and real-world datasets. Our results show that adversarial perturbations can significantly increase the FDR while maintaining high detection power, exposing fundamental limitations of current error-controlled novelty detection methods and motivating the development of more robust alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。