arXiv:2605.21780cs.LGcs.CR2026-05

通过隐私视角统一认证训练与推理阶段的后门攻击鲁棒性

Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy

论文配图:Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy
图 1 · 摘自论文原文
  • 引入差分隐私的对偶视角,将随机平滑扩展至联合扰动场景
  • 在MNIST和CIFAR-10上实现对训练/推理攻击的联合鲁棒性认证
  • 适用于复杂攻击模型,适合安全关键场景的模型验证

随机平滑是认证对抗扰动鲁棒性的有力工具,可分别用于训练时的随机化防御和测试时的随机化推断以应对投毒攻击与逃避攻击。然而,当训练与测试数据同时被扰动(即后门攻击)时,现有方法难以建立统一的鲁棒性证明,因需在同一证书中分析训练与测试阶段的随机机制。本文通过将随机平滑与差分隐私的对偶视角结合,利用隐私谱(privacy profiles)提供异构机制组合的数值化方法,构建了紧致、模块化、端到端的认证框架,可复用已有差分隐私机制的分析成果。我们以DP-SGD与深度分块聚合(Deep Partition Aggregation)结合推理时平滑的方式实例化该框架,实现了对训练时与推理时攻击的联合鲁棒性保证。在MNIST和CIFAR-10上的实验验证了该框架的有效性。总体而言,本工作提出了一种原则性强、通用性高的复合机制认证方法,能更好刻画现实攻击者的复杂能力。

原文摘要 · Abstract (English)

Randomized smoothing is a powerful tool for certifying robustness to adversarial perturbations, including poisoning attacks via randomized training and evasion attacks via randomized inference. Extending these guarantees to backdoor attacks, where training and test data are jointly perturbed, remains challenging because training- and test-time randomized mechanisms must be analyzed within a single robustness certificate. We address this by connecting randomized smoothing to the dual view of differential privacy through privacy profiles, which provide a numerical procedure for composing heterogeneous mechanisms. The resulting framework enables tight, modular, end-to-end certification of complex, composed mechanisms while leveraging existing analyses of differentially private mechanisms. We instantiate the framework for DP-SGD and Deep Partition Aggregation with inference-time smoothing, deriving joint robustness guarantees against both training-time and inference-time attacks. Experiments on MNIST and CIFAR-10 demonstrate the effectiveness of our framework. Overall, we provide a principled and general framework for using composite mechanisms to certify robustness under complex threat models that better capture the capabilities of real-world adversaries.

后门攻击差分隐私鲁棒性认证随机平滑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。