arXiv:2510.09288stat.MLcs.LG2025-10

用贝叶斯框架显式建模对抗攻击不确定性,提升模型鲁棒性。

A unifying Bayesian framework for adversarial robustness

  • 构建贝叶斯框架,将对抗攻击视为随机信道,明确概率假设。
  • 提出训练时的主动防御与推理时的被动净化两种策略。
  • 统一多种先进防御方法,实证显示显式建模更有效。

机器学习模型对对抗攻击的脆弱性仍是重大社会安全挑战。传统防御如对抗训练通常通过最小化最坏情况损失来增强鲁棒性,但这些确定性方法未考虑对手攻击中的不确定性。虽存在随机防御方法为对手攻击分配概率分布,但往往缺乏统计严谨性,且未明示其假设。为此,我们提出一个形式化的贝叶斯框架,通过随机信道建模对抗不确定性,并清晰阐述所有概率假设。该框架导出两种鲁棒化策略:训练阶段的主动防御(与对抗训练一致),以及运行阶段的被动防御(与对抗净化一致)。若干前沿防御方法可作为本模型的极限情形被恢复。我们通过实验验证了该方法的有效性,展示了显式建模对抗不确定性的优势。

原文摘要 · Abstract (English)

The vulnerability of machine learning models to adversarial attacks remains a critical societal security challenge. Traditional defenses, such as adversarial training, typically robustify models by minimizing a worst-case loss. These deterministic approaches do not account for uncertainty in the adversary's attack. While stochastic defenses placing a probability distribution on the adversary exist, they often lack statistical rigor and fail to make explicit their underlying assumptions. To resolve these issues, we introduce a formal Bayesian framework that models adversarial uncertainty through a stochastic channel, articulating all probabilistic assumptions. This yields two robustification strategies: a proactive defense enacted during training, aligned with adversarial training, and a reactive defense enacted during operations, aligned with adversarial purification. Several state-of-the-art defenses can be recovered as limiting cases of our model. We empirically validate our methodology, showcasing the benefits of explicitly modeling adversarial uncertainty.

贝叶斯对抗鲁棒性不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。