arXiv:2409.19431stat.MLcs.IT2024-09ICML被引 5

提出负倾斜经验风险的泛化与鲁棒性理论,为无分布偏移下的机器学习提供新工具。

Generalization and Robustness of the Tilted Empirical Risk

  • 基于负倾斜经验风险,推导出统一的信息论泛化界。
  • 在损失函数矩有界条件下,收敛速率达 $O(n^{-ε/(1+ε)})$。
  • 适用于噪声数据和分布偏移场景,支持数据驱动选择倾斜参数。

监督统计学习算法的泛化误差(风险)衡量其对未见数据的预测能力。受指数倾斜启发, extcite{li2020tilted} 提出倾斜经验风险(TER)作为分类与回归等任务的非线性风险度量。本文研究负倾斜下 TER 的泛化误差,在损失函数无界但 $(1+ε)$-阶矩有界的条件下,给出统一的信息论泛化误差上界,收敛率为 $O(n^{-ε/(1+ε)})$,揭示了 TER 在无分布偏移下的新应用。其次,分析了训练时噪声异常值对 TER 的鲁棒性,并在分布偏移下提供理论保证。通过简单实验验证理论结果,实现基于边界估计的数据驱动倾斜参数选择。

原文摘要 · Abstract (English)

The generalization error (risk) of a supervised statistical learning algorithm quantifies its prediction ability on previously unseen data. Inspired by exponential tilting, \citet{li2020tilted} proposed the {\it tilted empirical risk} (TER) as a non-linear risk metric for machine learning applications such as classification and regression problems. In this work, we examine the generalization error of the tilted empirical risk in the robustness regime under \textit{negative tilt}. Our first contribution is to provide uniform and information-theoretic bounds on the {\it tilted generalization error}, defined as the difference between the population risk and the tilted empirical risk, under negative tilt for unbounded loss function under bounded $(1+ε)$-th moment of loss function for some $ε\in(0,1]$ with a convergence rate of $O(n^{-ε/(1+ε)})$ where $n$ is the number of training samples, revealing a novel application for TER under no distribution shift. Secondly, we study the robustness of the tilted empirical risk with respect to noisy outliers at training time and provide theoretical guarantees under distribution shift for the tilted empirical risk. We empirically corroborate our findings in simple experimental setups where we evaluate our bounds to select the value of tilt in a data-driven manner.

泛化误差鲁棒性经验风险倾斜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。