用β散度训练神经网络,显著提升对异常值的鲁棒性。
Provably robust learning of regression neural networks using $β$-divergences
- 基于β-散度设计新框架rRNet,适配多种非光滑模型
- 理论证明可达到50%渐近崩溃点,抗干扰能力更强
- 适合数据含噪声或异常值的回归任务,如真实观测场景
回归神经网络通常通过最小化均方误差训练,对异常值和数据污染敏感。现有鲁棒训练方法多局限于特定场景,且缺乏充分理论支撑。本文提出基于β-散度(密度幂散度)的新鲁棒学习框架rRNet,适用于包括非光滑激活函数和误差分布在内的广泛回归神经网络,并可退化为经典最大似然估计。rRNet采用交替优化实现,其在温和可验证条件下收敛至驻点。参数估计与预测器的影响函数被理论刻画,当β∈(0,1]时,影响函数有界。进一步证明rRNet在假设模型下具有最优50%渐近崩溃点,提供强全局鲁棒性保障。理论结果通过模拟实验与真实数据分析验证,表明rRNet在函数逼近与带噪声预测任务中优于现有方法。
原文摘要 · Abstract (English)
Regression neural networks (NNs) are most commonly trained by minimizing the mean squared prediction error, which is highly sensitive to outliers and data contamination. Existing robust training methods for regression NNs are often limited in scope and rely primarily on empirical validation, with only a few offering partial theoretical guarantees. In this paper, we propose a new robust learning framework for regression NNs based on the $β$-divergence (also known as the density power divergence) which we call `rRNet'. It applies to a broad class of regression NNs, including models with non-smooth activation functions and error densities, and recovers the classical maximum likelihood learning as a special case. The rRNet is implemented via an alternating optimization scheme, for which we establish convergence guarantees to stationary points under mild, verifiable conditions. The (local) robustness of rRNet is theoretically characterized through the influence functions of both the parameter estimates and the resulting rRNet predictor, which are shown to be bounded for suitable choices of the tuning parameter $β$, depending on the error density. We further prove that rRNet attains the optimal 50\% asymptotic breakdown point at the assumed model for all $β\in(0, 1]$, providing a strong global robustness guarantee that is largely absent for existing NN learning methods. Our theoretical results are complemented by simulation experiments and real-data analyses, illustrating practical advantages of rRNet over existing approaches in both function approximation problems and prediction tasks with noisy observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。