arXiv:2603.17628stat.MLcs.AI2026-03被引 1

统一应对标签噪声与对抗攻击的鲁棒神经网络训练方法

rSDNet: Unified Robust Neural Learning against Label Noise and Adversarial Attacks

  • 基于S-散度的统一损失函数,自动降低异常样本影响
  • 在三个图像数据集上同时提升对标签噪声和对抗攻击的鲁棒性
  • 适合需要高可靠性的实际部署场景,如医疗或自动驾驶

神经网络在现代人工智能中至关重要,但其训练对数据污染极为敏感。标准分类器通过最小化类别交叉熵损失进行训练,该方法在理想条件下统计效率高,但易受标签噪声(输出空间污染)和对抗扰动(输入空间最坏偏离)影响。本文提出rSDNet,一种基于S-散度的统一鲁棒学习框架,将神经网络训练建模为最小散度估计问题。该方法继承经典统计估计的鲁棒性,通过模型概率自动降权异常样本。我们建立了rSDNet的群体性质:费希尔一致性、分类校准(蕴含贝叶斯最优性),以及在均匀标签噪声和微小特征污染下的鲁棒性保证。在三个基准图像分类数据集上的实验表明,rSDNet在保持干净数据上竞争力准确率的同时,显著提升了对标签污染和对抗攻击的鲁棒性。结果凸显最小散度学习是应对异质数据污染的原理性有效框架。

原文摘要 · Abstract (English)

Neural networks are central to modern artificial intelligence, yet their training remains highly sensitive to data contamination. Standard neural classifiers are trained by minimizing the categorical cross-entropy loss, corresponding to maximum likelihood estimation under a multinomial model. While statistically efficient under ideal conditions, this approach is highly vulnerable to contaminated observations including label noises corrupting supervision in the output space, and adversarial perturbations inducing worst-case deviations in the input space. In this paper, we propose a unified and statistically grounded framework for robust neural classification that addresses both forms of contamination within a single learning objective. We formulate neural network training as a minimum-divergence estimation problem and introduce rSDNet, a robust learning algorithm based on the general class of $S$-divergences. The resulting training objective inherits robustness properties from classical statistical estimation, automatically down-weighting aberrant observations through model probabilities. We establish essential population-level properties of rSDNet, including Fisher consistency, classification calibration implying Bayes optimality, and robustness guarantees under uniform label noise and infinitesimal feature contamination. Experiments on three benchmark image classification datasets show that rSDNet improves robustness to label corruption and adversarial attacks while maintaining competitive accuracy on clean data, Our results highlight minimum-divergence learning as a principled and effective framework for robust neural classification under heterogeneous data contamination.

鲁棒学习标签噪声对抗攻击S-散度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。