arXiv:2505.07367cs.LGmath.ST2025-05被引 6

提出自选择数据学习的泛化边界与停止规则,无需假设数据分布。

Generalization Bounds and Stopping Rules for Learning with Self-Selected Data

  • 基于覆盖数与Wasserstein模糊集,建立统一泛化界。
  • 无需数据分布假设,适用于收敛与有限迭代解。
  • 给出可验证的停止规则,适合关注泛化性能的实践者。

许多学习范式根据先前学习的参数自我选择训练数据,如主动学习、半监督学习、强化学习和提升方法。Rodemann等(2024)将其统一为“互惠学习”框架。本文研究此类方法从自选样本中泛化的能力。我们利用覆盖数与Wasserstein模糊集,证明了互惠学习的普适泛化界,不依赖于自选数据的分布假设,仅需算法层面的可验证条件。结果涵盖收敛解与有限迭代解;后者具有任意时刻有效性,从而为追求模型外样本性能保障的实践者提供停止规则。最后,我们在互惠学习的特例——半监督学习中展示了我们的边界与停止规则。

原文摘要 · Abstract (English)

Many learning paradigms self-select training data in light of previously learned parameters. Examples include active learning, semi-supervised learning, bandits, or boosting. Rodemann et al. (2024) unify them under the framework of "reciprocal learning". In this article, we address the question of how well these methods can generalize from their self-selected samples. In particular, we prove universal generalization bounds for reciprocal learning using covering numbers and Wasserstein ambiguity sets. Our results require no assumptions on the distribution of self-selected data, only verifiable conditions on the algorithms. We prove results for both convergent and finite iteration solutions. The latter are anytime valid, thereby giving rise to stopping rules for a practitioner seeking to guarantee the out-of-sample performance of their reciprocal learning algorithm. Finally, we illustrate our bounds and stopping rules for reciprocal learning's special case of semi-supervised learning.

泛化理论自选择数据停止规则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。