证明正则化最小二乘在重尾噪声下仍能实现最优收敛速度。
Regularized least squares learning with heavy-tailed noise is minimax optimal
- 基于积分算子框架,结合希尔伯特空间的弗克-那加夫不等式推导风险界。
- 在仅存在有限高阶矩的重尾噪声下,达到与亚指数噪声相同的最优收敛率。
- 适用于对鲁棒性要求高的机器学习场景,如金融或传感器数据建模。
本文研究了在噪声仅有有限高阶矩的情况下,再生核希尔伯特空间中岭回归的性能。通过经典的积分算子框架,建立了包含次高斯项和多项式项的过量风险界。其中主导的次高斯成分使得收敛速率可达到以往仅在亚指数噪声假设下才成立的水平。该速率在标准特征值衰减条件下为最优,证明了正则化最小二乘对重尾噪声具有渐近鲁棒性。推导过程依赖于希尔伯特空间取值随机变量的弗克-那加夫不等式。
原文摘要 · Abstract (English)
This paper examines the performance of ridge regression in reproducing kernel Hilbert spaces in the presence of noise that exhibits a finite number of higher moments. We establish excess risk bounds consisting of subgaussian and polynomial terms based on the well known integral operator framework. The dominant subgaussian component allows to achieve convergence rates that have previously only been derived under subexponential noise - a prevalent assumption in related work from the last two decades. These rates are optimal under standard eigenvalue decay conditions, demonstrating the asymptotic robustness of regularized least squares against heavy-tailed noise. Our derivations are based on a Fuk-Nagaev inequality for Hilbert-space valued random variables.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。