arXiv:2509.01235cs.LGcond-mat.stat-mech2025-09

通过几何感知训练提升模型抗干扰能力,让特征更紧凑分离。

Geometric origin of adversarial vulnerability in deep learning

  • 逐层局部训练塑造特征空间几何结构
  • 实现类内紧凑、类间分离,增强对抗鲁棒性
  • 适合关注模型安全性与表征学习的研究者

平衡训练准确率与对抗鲁棒性一直是深度学习的核心挑战。本文提出一种几何感知的深度学习框架,通过逐层局部训练来塑造深层神经网络的内部表示。该方法在特征空间中促进类内紧凑性和类间分离性,从而实现流形平滑,并有效抵御白盒或黑盒攻击。性能可由lue{参数积分后的数据依赖统计力学}解释,辅以包含隐层表示元素间赫布耦合的拟经验模型。基于此几何感知学习框架,深度网络可在不增加表示干扰的前提下,将新信息融入已有知识结构。

原文摘要 · Abstract (English)

Balancing training accuracy and adversarial robustness has beeen a challenge since the birth of deep learning. Here, we introduce a geometry-aware deep learning framework that leverages layer-wise local training to sculpt the internal representations of deep neural networks. This framework promotes intra-class compactness and inter-class separation in feature space, leading to manifold smoothness and adversarial robustness against white or black box attacks. The performance can be explained by \blue{data-dependent statistical mechanics of integrating out the network parameters}, \blue{supplemented by a phenomenological model} with Hebbian coupling between elements of the hidden representation. Based on the current geometry-aware learning framework, the deep network can assimilate new information into existing knowledge structures while reducing representation interference.

对抗样本特征空间深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。