arXiv:2511.14003cs.LGcs.CR2025-11

攻击者可伪造证书让模型误判却仍获得虚假的鲁棒性证明。

Certified but Fooled! Breaking Certified Defences with Ghost Certificates

  • 设计聚焦区域的对抗样本,伪装成合法认证输入。
  • 在ImageNet上成功绕过Densepure等顶尖认证防御,扰动极小且不可见。
  • 揭示现有认证机制漏洞,适合安全研究者参考。

认证防御承诺可证明的鲁棒性保证。本文研究恶意利用概率认证框架的潜在风险,目标不仅是误导分类器,更是在认证过程中伪造对抗样本的鲁棒性证明。近期ICLR研究发现,通过构造大扰动可使输入进入能生成错误类别证书的区域。本文进一步探究:是否可通过微小、不可察觉的扰动,既导致误分类,又诱使认证模型为目标类别生成远大于源类别的鲁棒半径?我们提出区域聚焦的对抗样本方法,实现证书欺骗,并在ImageNet上验证其有效性,成功绕过Densepure等前沿认证防御。结果表明当前认证机制存在严重局限,亟需重新审视其理论边界。

原文摘要 · Abstract (English)

Certified defenses promise provable robustness guarantees. We study the malicious exploitation of probabilistic certification frameworks to better understand the limits of guarantee provisions. Now, the objective is to not only mislead a classifier, but also manipulate the certification process to generate a robustness guarantee for an adversarial input certificate spoofing. A recent study in ICLR demonstrated that crafting large perturbations can shift inputs far into regions capable of generating a certificate for an incorrect class. Our study investigates if perturbations needed to cause a misclassification and yet coax a certified model into issuing a deceptive, large robustness radius for a target class can still be made small and imperceptible. We explore the idea of region-focused adversarial examples to craft imperceptible perturbations, spoof certificates and achieve certification radii larger than the source class ghost certificates. Extensive evaluations with the ImageNet demonstrate the ability to effectively bypass state-of-the-art certified defenses such as Densepure. Our work underscores the need to better understand the limits of robustness certification methods.

认证防御对抗攻击鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。