提出无需参数可识别的鲁棒神经分类器一致性理论,解决噪声数据下的训练难题。
No Unique Minimizer, No Problem: On the Consistency of Robust Neural Classifiers
- 基于S散度族设计鲁棒训练方法,不依赖参数可识别性
- 证明经验最小化解收敛到最优等价类,且在三类架构中成立
- 适用于对噪声和对抗样本敏感的视觉与语言任务
通过交叉熵最小化训练的神经网络分类器对标签噪声和对抗污染极为敏感。尽管鲁棒方法能提供有界影响并抵抗干扰,但在深度学习场景下其统计基础不足,根源在于神经参数化不可识别,导致总体损失最小化器是一个参数等价类而非唯一点。本文基于S散度族发展了一种无需可识别性假设的一致性理论。将训练视为非可识别参数空间上的随机优化问题,证明在温和正则条件下,经验S散度最小化器会收敛到总体最优等价类,并验证了三种架构选择均满足该条件。进一步证明鲁棒训练算法的极限点是经验目标函数的驻点。在视觉与语言基准数据集上的实验表明,S散度训练在保持干净数据准确率的同时,性能与现有鲁棒方法相当。
原文摘要 · Abstract (English)
Neural network classifiers trained by cross-entropy minimization are highly sensitive to label noise and adversarial contamination. While robust alternatives offer bounded influence and resistance to corruption, their statistical foundations in the deep learning setting are insufficient due to a fundamental difficulty: neural parameterizations are non-identifiable, so the population loss minimizer is an equivalence class of parameters, not a unique point. We develop a consistency theory for robust neural classifiers based on the S-divergence family that requires no identifiability assumption. Casting training as stochastic optimization over a non-identifiable parameter space, we prove that empirical S-divergence minimizers converge to the population-optimal equivalence class under mild regularity conditions, and verify these conditions for three architecture choices. We further establish that limit points of the robust training algorithm are stationary points of the empirical objective. Experiments on vision and language benchmark datasets confirm that S-divergence training maintains clean-data accuracy while exhibiting performance competitive with existing robust methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。