arXiv:2509.00089cs.LG2025-09

让多个模型协作对抗攻击,提升防御能力。

Learning from Peers: Collaborative Ensemble Adversarial Training

  • 通过预测差异动态加权样本,促进模型间协作学习。
  • 在CIFAR-10、ImageNet等数据集上超越现有方法。
  • 不依赖特定模型,可通用适配各类集成框架。

集成对抗训练(EAT)通过多模型协同提升模型对对抗攻击的鲁棒性,但现有方法多独立训练子模型,忽视了模型间的合作潜力。我们发现,子模型预测差异较大的样本更接近集成决策边界,对整体鲁棒性影响更大。为此,提出高效且新颖的协作式集成对抗训练(CEAT),在对抗训练中赋予子模型间预测差异大的样本更高权重。通过引入校准距离正则化,利用概率差异自适应分配样本权重。在广泛采用的CIFAR-10、ImageNet等数据集上的大量实验表明,所提方法在鲁棒性上达到当前最优性能。尤为关键的是,CEAT为模型无关设计,可无缝融入多种集成方法,具备高度灵活性。

原文摘要 · Abstract (English)

Ensemble Adversarial Training (EAT) attempts to enhance the robustness of models against adversarial attacks by leveraging multiple models. However, current EAT strategies tend to train the sub-models independently, ignoring the cooperative benefits between sub-models. Through detailed inspections of the process of EAT, we find that that samples with classification disparities between sub-models are close to the decision boundary of ensemble, exerting greater influence on the robustness of ensemble. To this end, we propose a novel yet efficient Collaborative Ensemble Adversarial Training (CEAT), to highlight the cooperative learning among sub-models in the ensemble. To be specific, samples with larger predictive disparities between the sub-models will receive greater attention during the adversarial training of the other sub-models. CEAT leverages the probability disparities to adaptively assign weights to different samples, by incorporating a calibrating distance regularization. Extensive experiments on widely-adopted datasets show that our proposed method achieves the state-of-the-art performance over competitive EAT methods. It is noteworthy that CEAT is model-agnostic, which can be seamlessly adapted into various ensemble methods with flexible applicability.

对抗训练模型集成鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。