解析经验风险最小化的高概率误差界,给出可复用的推导框架。
A Researcher's Guide to Empirical Risk Minimization
- 基于三步法:基本不等式、局部一致集中、不动点论证,统一推导误差界。
- 在温和方差-风险条件下,用局部Rademacher复杂度定义临界半径并给出具体上界。
- 适用于因果推断、缺失数据等带干扰项的问题,支持样本分裂与原样本场景。
本指南为经验风险最小化(ERM)提供高概率后悔率界限参考。内容模块化:先阐述直观思想和通用证明策略,再在高层假设下给出广泛适用的保证,并提供验证特定损失与函数类条件的工具。强调多数ERM收敛速率可归纳为三步法——基本不等式、一致局部集中界、不动点论证——在满足温和的Bernstein型方差-风险条件下,得到以局部Rademacher复杂度定义的临界半径为表达形式的后悔界。为使边界具体化,利用局部最大不等式与度量熵积分对临界半径进行上界估计,从而恢复了VC-子图、Sobolev/Holder及有界变差类的经典收敛速率。同时研究含干扰项的ERM,包括加权ERM与Neyman正交损失,这些在因果推断、缺失数据与域适应中常见。依据正交统计学习框架,指出此类问题常具备后悔转移界,将估计损失下的后悔分解为(1)估计损失下的统计误差与(2)干扰估计带来的近似误差。在样本分割或交叉拟合下,第一项可用标准固定损失的ERM后悔界控制,第二项仅依赖于干扰估计精度。作为新贡献,还处理了原样本情形,即干扰项与ERM在同一数据上拟合,推导出后悔界,并表明在适当光滑性与Donsker型条件下仍可达到快速最优率。
原文摘要 · Abstract (English)
This guide provides a reference for high-probability regret bounds in empirical risk minimization (ERM). The presentation is modular: we begin with intuition and general proof strategies, then state broadly applicable guarantees under high-level conditions and provide tools for verifying them for specific losses and function classes. We emphasize that many ERM rate derivations can be organized around a three-step recipe -- a basic inequality, a uniform local concentration bound, and a fixed-point argument -- which yields regret bounds in terms of a critical radius, defined via localized Rademacher complexity, under a mild Bernstein-type variance-risk condition. To make these bounds concrete, we upper bound the critical radius using local maximal inequalities and metric-entropy integrals, thereby recovering familiar rates for VC-subgraph, Sobolev/Hölder, and bounded-variation classes. We also study ERM with nuisance components -- including weighted ERM and Neyman-orthogonal losses -- as they arise in causal inference, missing data, and domain adaptation. Following the orthogonal statistical learning framework, we highlight that these problems often admit regret-transfer bounds linking regret under an estimated loss to population regret under the target loss. These bounds typically decompose the regret into (i) statistical error under the estimated loss and (ii) approximation error due to nuisance estimation. Under sample splitting or cross-fitting, the first term can be controlled using standard fixed-loss ERM regret bounds, while the second depends only on nuisance-estimation accuracy. As a novel contribution, we also treat the in-sample regime, in which the nuisances and the ERM are fit on the same data, deriving regret bounds and showing that fast oracle rates remain attainable under suitable smoothness and Donsker-type conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。