arXiv:2512.22880cs.LGstat.ML2025-12被引 15

提出新型一致性理论,精准评估机器学习中代理损失与真实损失的差距。

Fundamental Novel Consistency Theory: $H$-Consistency Bounds

  • 基于假设集构建 $H$-一致性边界,比传统方法更严格可靠。
  • 首次为多分类任务中的最大、求和及约束损失提供一致性边界。
  • 揭示某些情况下无法获得非平凡边界,指导模型与损失函数选择。

在机器学习中,训练时优化的损失函数常因计算不可行或不可微而与任务性能定义的目标损失不同。本文深入研究目标损失估计误差相对于代理损失估计误差的关系,提出考虑假设集 $H$ 的 $H$-一致性边界。这些边界强于贝叶斯一致性或 $H$-校准,且比过量误差界更具信息量。我们从二分类出发,建立紧致的分布相关与无关边界;针对凸代理损失(包括线性模型和神经网络)给出显式边界,并分析 $ρ$-间隔与逻辑斯蒂损失等在对抗情形下的表现。扩展至多分类任务,首次为最大、求和及约束损失提供 $H$-一致性边界,覆盖非对抗与对抗场景。证明在某些情况下,非平凡 $H$-一致性边界不可达。同时研究了 comp-sum 损失(如交叉熵、MAE),首次推导其 $H$-一致性边界,并引入平滑对抗变体以实现鲁棒学习。本文建立统一框架,为多种代理损失推导边界,提出对约束与 comp-sum 损失的新表征。最后分析 $H$-一致性边界的增长速率,证明光滑代理损失在二分类与多分类任务中具有普适的平方根增长规律,并通过最小化间隙分析指导代理损失选择。

原文摘要 · Abstract (English)

In machine learning, the loss functions optimized during training often differ from the target loss that defines task performance due to computational intractability or lack of differentiability. We present an in-depth study of the target loss estimation error relative to the surrogate loss estimation error. Our analysis leads to $H$-consistency bounds, which are guarantees accounting for the hypothesis set $H$. These bounds offer stronger guarantees than Bayes-consistency or $H$-calibration and are more informative than excess error bounds. We begin with binary classification, establishing tight distribution-dependent and -independent bounds. We provide explicit bounds for convex surrogates (including linear models and neural networks) and analyze the adversarial setting for surrogates like $ρ$-margin and sigmoid loss. Extending to multi-class classification, we present the first $H$-consistency bounds for max, sum, and constrained losses, covering both non-adversarial and adversarial scenarios. We demonstrate that in some cases, non-trivial $H$-consistency bounds are unattainable. We also investigate comp-sum losses (e.g., cross-entropy, MAE), deriving their first $H$-consistency bounds and introducing smooth adversarial variants that yield robust learning algorithms. We develop a comprehensive framework for deriving these bounds across various surrogates, introducing new characterizations for constrained and comp-sum losses. Finally, we examine the growth rates of $H$-consistency bounds, establishing a universal square-root growth rate for smooth surrogates in binary and multi-class tasks, and analyze minimizability gaps to guide surrogate selection.

理论分析损失函数一致性多分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。