arXiv:2503.07818cs.LG2025-03

提升神经网络内部抗干扰能力,让模型更稳更准。

Strengthening the Internal Adversarial Robustness in Lifted Neural Networks

  • 通过改进训练损失函数,增强网络内部对对抗攻击的防御能力。
  • 新方法融合有目标与无目标扰动,使模型在多种攻击下仍保持稳定。
  • 适合关注模型鲁棒性、安全性与泛化性能的研究者。

提升神经网络(Lifted Neural Networks)中内部活动的对抗鲁棒性,可通过结合特定类型的对抗训练实现,不仅增强输入层和内部层的鲁棒性,还能提升泛化性能。本文首先研究仅通过修改训练损失函数,能否进一步加强该框架下的对抗鲁棒性。随后,针对现有方法的若干局限性进行修正,提出一种新型训练损失函数,同时整合有目标与无目标对抗扰动,显著提升模型在复杂攻击场景下的稳定性与可靠性。

原文摘要 · Abstract (English)

Lifted neural networks (i.e. neural architectures explicitly optimizing over respective network potentials to determine the neural activities) can be combined with a type of adversarial training to gain robustness for internal as well as input layers, in addition to improved generalization performance. In this work we first investigate how adversarial robustness in this framework can be further strengthened by solely modifying the training loss. In a second step we fix some remaining limitations and arrive at a novel training loss for lifted neural networks, that combines targeted and untargeted adversarial perturbations.

神经网络对抗鲁棒性训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。