提出新方法在重尾噪声下实现非参数回归的稳健学习,突破传统理论局限。
Understanding Robust Machine Learning for Nonparametric Regression with Heavy-Tailed Noise
- 引入概率有效假设空间,解决无界函数带来的分析难题。
- 建立鲁棒风险与预测误差的关联,揭示参数σ的鲁棒性-偏差权衡。
- 给出无需有界假设的有限样本误差界,适用于多种鲁棒损失。
研究重尾噪声下的鲁棒非参数回归,其中假设类可能包含无界函数,通过鲁棒损失函数ℓ_σ确保鲁棒性。以再生核希尔伯特空间(RKHS)中Tikhonov正则化风险最小化下的Huber回归为例,解决两大核心挑战:(i) 弱矩条件下标准浓度工具失效;(ii) 无界假设空间带来的分析困难。首先提出概念性观点:传统鲁棒损失的泛化误差界无法准确反映外样本性能,应以预测误差(即到真值f⋆的L₂距离)作为可学习性度量,该误差与σ无关,直接对应鲁棒估计目标。为在无界性下可操作,提出“概率有效假设空间”,使估计器在高概率下受限,支持在弱(1+ε)阶矩条件下进行有意义的偏差-方差分解。技术上,建立新比较定理,将超额鲁棒风险与L₂预测误差关联,残差阶为O(σ^{-2ε}),阐明尺度参数σ引发的鲁棒性-偏差权衡。基于此,推导出在无统一有界性和重尾噪声下,对RKHS中Huber回归的显式有限样本误差界与收敛率。研究提供合理调参规则,可推广至其他鲁棒损失,并强调预测误差而非超额泛化风险,才是分析鲁棒学习的根本视角。
原文摘要 · Abstract (English)
We investigate robust nonparametric regression in the presence of heavy-tailed noise, where the hypothesis class may contain unbounded functions and robustness is ensured via a robust loss function $\ell_σ$. Using Huber regression as a close-up example within Tikhonov-regularized risk minimization in reproducing kernel Hilbert spaces (RKHS), we address two central challenges: (i) the breakdown of standard concentration tools under weak moment assumptions, and (ii) the analytical difficulties introduced by unbounded hypothesis spaces. Our first message is conceptual: conventional generalization-error bounds for robust losses do not faithfully capture out-of-sample performance. We argue that learnability should instead be quantified through prediction error, namely the $L_2$-distance to the truth $f^\star$, which is $σ$-independent and directly reflects the target of robust estimation. To make this workable under unboundedness, we introduce a \emph{probabilistic effective hypothesis space} that confines the estimator with high probability and enables a meaningful bias--variance decomposition under weak $(1+ε)$-moment conditions. Technically, we establish new comparison theorems linking the excess robust risk to the $L_2$ prediction error up to a residual of order $\mathcal{O}(σ^{-2ε})$, clarifying the robustness--bias trade-off induced by the scale parameter $σ$. Building on this, we derive explicit finite-sample error bounds and convergence rates for Huber regression in RKHS that hold without uniform boundedness and under heavy-tailed noise. Our study delivers principled tuning rules, extends beyond Huber to other robust losses, and highlights prediction error, not excess generalization risk, as the fundamental lens for analyzing robust learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。