arXiv:2501.18581cs.LG2025-01JMLR被引 7

只有特定损失函数能实现清晰的偏差-方差分解,其他都不行。

Bias-variance decompositions: the exclusive privilege of Bregman divergences

  • 证明了仅g-Bregman或rho-tau散度可实现清晰分解
  • 0-1损失和L1损失无法做到,解释了此前失败原因
  • 适合研究模型泛化理论的学者参考

偏差-方差分解广泛用于理解机器学习模型的泛化性能。虽然平方误差损失可直接分解,但零一损失或L1损失要么无法使偏差与方差之和等于期望损失,要么依赖缺乏合理性质的定义。近期研究表明,更广泛的Bregman散度类可实现清晰分解,交叉熵是特例。然而,这些分解成立的充要条件仍不清楚。本文在连续、非负且满足不可区分性恒等(两参数相同时损失为零)的弱正则条件下,研究损失函数,证明g-Bregman或rho-tau散度是唯一具备清晰偏差-方差分解的损失函数。g-Bregman散度可通过可逆变量变换转化为标准Bregman散度。因此,经变量变换后,平方马氏距离是唯一具有清晰分解的对称损失函数。常见度量如0-1和L1损失无法实现清晰分解,解释了以往尝试失败的原因。我们还分析了放宽损失函数限制的影响及其对结果的改变。

原文摘要 · Abstract (English)

Bias-variance decompositions are widely used to understand the generalization performance of machine learning models. While the squared error loss permits a straightforward decomposition, other loss functions - such as zero-one loss or $L_1$ loss - either fail to sum bias and variance to the expected loss or rely on definitions that lack the essential properties of meaningful bias and variance. Recent research has shown that clean decompositions can be achieved for the broader class of Bregman divergences, with the cross-entropy loss as a special case. However, the necessary and sufficient conditions for these decompositions remain an open question. In this paper, we address this question by studying continuous, nonnegative loss functions that satisfy the identity of indiscernibles (zero loss if and only if the two arguments are identical), under mild regularity conditions. We prove that so-called $g$-Bregman or rho-tau divergences are the only such loss functions that have a clean bias-variance decomposition. A $g$-Bregman divergence can be transformed into a standard Bregman divergence through an invertible change of variables. This makes the squared Mahalanobis distance, up to such a variable transformation, the only symmetric loss function with a clean bias-variance decomposition. Consequently, common metrics such as $0$-$1$ and $L_1$ losses cannot admit a clean bias-variance decomposition, explaining why previous attempts have failed. We also examine the impact of relaxing the restrictions on the loss functions and how this affects our results.

泛化分析损失函数理论研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。