提出一种新方法优化风险最小化问题,提升泛化能力。
Empirical Risk Minimization with $f$-Divergence Regularization
- 引入归一化函数解决 $f$-散度正则化下的经验风险最小化问题。
- 证明解的理论性质,并给出数值算法逼近归一化因子。
- 揭示不同散度函数对模型训练与测试风险的影响,适合理论研究者。
本文提出经验风险最小化结合 $f$-散度正则化(ERM-$f$DR)的求解方法,并建立其等价于在 $f$-散度约束下最小化期望经验风险的条件。该方法拓展了适用的 $f$-散度范围,恢复了已有结果。核心贡献是引入归一化函数,通过非线性常微分方程(ODE)隐式刻画其性质,并据此设计数值算法,在温和假设下近似计算归一化因子。进一步分析表明,不同 $f$-散度的 ERM-$f$DR 问题可通过经验风险变换实现结构等价。数值实验展示了在不同 $f$-散度正则化下训练与测试风险的变化,凸显选择 $f$ 函数的实际影响。
原文摘要 · Abstract (English)
In this paper, the solution to the empirical risk minimization problem with $f$-divergence regularization (ERM-$f$DR) is presented and conditions under which the solution also serves as the solution to the minimization of the expected empirical risk subject to an $f$-divergence constraint are established. The proposed approach extends applicability to a broader class of $f$-divergences than previously reported and yields theoretical results that recover previously known results. Additionally, the difference between the expected empirical risk of the ERM-$f$DR solution and that of its reference measure is characterized, providing insights into previously studied cases of $f$-divergences. A central contribution is the introduction of the normalization function, a mathematical object that is critical in both the dual formulation and practical computation of the ERM-$f$DR solution. This work presents an implicit characterization of the normalization function as a nonlinear ordinary differential equation (ODE), establishes its key properties, and subsequently leverages them to construct a numerical algorithm for approximating the normalization factor under mild assumptions. Further analysis demonstrates structural equivalences between ERM-$f$DR problems with different $f$-divergences via transformations of the empirical risk. Finally, the proposed algorithm is used to compute the training and test risks of ERM-$f$DR solutions under different $f$-divergence regularizers. This numerical example highlights the practical implications of choosing different functions $f$ in ERM-$f$DR problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。