arXiv:2506.12829stat.MLcs.LG2025-06

提出可估计的统一学习界,解决分布偏移下的泛化难题

General and Estimable Learning Bound Unifying Covariate and Concept Shifts

  • 用熵正则最优传输定义新偏移度量,摆脱支持集匹配限制
  • 导出适用于广义损失与随机标签的统一误差上界
  • 开发可估的DataShifts算法,实现在多数场景中量化偏移

在分布偏移下的泛化仍是现代机器学习的核心挑战,但现有学习界理论局限于狭窄的理想化设定且无法从样本中估计。本文弥合理论与实际应用之间的鸿沟。我们首先指出,现有界在源与目标支持集不匹配时失效,因其概念偏移定义崩溃。基于熵正则最优传输,我们提出对支持集无关的协变量与概念偏移新定义,并推导出适用于广泛损失函数、标签空间及随机标注的新型统一误差界。进一步,我们开发了这些偏移的可估计器并提供集中率保证,以及DataShifts算法——可在大多数应用场景中量化分布偏移并估计误差界,为分布偏移下的学习误差分析提供严格且通用的工具。

原文摘要 · Abstract (English)

Generalization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-estimable from samples. In this paper, we bridge the gap between theory and practical applications. We first show that existing bounds become loose and non-estimable because their concept shift definition breaks when the source and target supports mismatch. Leveraging entropic optimal transport, we propose new support-agnostic definitions for covariate and concept shifts, and derive a novel unified error bound that applies to broad loss functions, label spaces, and stochastic labeling. We further develop estimators for these shifts with concentration guarantees, and the DataShifts algorithm, which can quantify distribution shifts and estimate the error bound in most applications -- a rigorous and general tool for analyzing learning error under distribution shift.

学习界分布偏移可估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。