arXiv:2410.22887stat.MLcs.IT2024-10NeurIPS被引 4

用条件f-信息推导新泛化界,不依赖传统上界技巧。

Generalization Bounds via Conditional $f$-Information

  • 基于条件f-信息框架,无需上界控制累积生成函数。
  • 在超样本设定下适用于有界与无界损失函数。
  • 实证显示新界优于旧界,适合理论研究者参考。

本文提出一种基于条件f-信息的新信息论泛化界,扩展了传统的条件互信息框架。我们给出一种通用方法,在超样本设定下通过f-信息推导泛化界,适用于有界和无界损失函数。不同于以往基于互信息的界,我们的证明策略不依赖于对互信息变分公式中累积生成函数的上界估计,而是通过精心选择变分公式中的可测函数,将累积生成函数或其上界设为零。尽管部分技术受近期硬币赌博框架(如Jang等,2023)启发,但结果独立于在线赌博算法的后悔保证。此外,我们新推导的互信息界可恢复许多已有结果,并深化了对其潜在局限性的理解。最后,我们对多种f-信息度量进行了泛化性能的实证比较,验证了新界的优越性。

原文摘要 · Abstract (English)

In this work, we introduce novel information-theoretic generalization bounds using the conditional $f$-information framework, an extension of the traditional conditional mutual information (MI) framework. We provide a generic approach to derive generalization bounds via $f$-information in the supersample setting, applicable to both bounded and unbounded loss functions. Unlike previous MI-based bounds, our proof strategy does not rely on upper bounding the cumulant-generating function (CGF) in the variational formula of MI. Instead, we set the CGF or its upper bound to zero by carefully selecting the measurable function invoked in the variational formula. Although some of our techniques are partially inspired by recent advances in the coin-betting framework (e.g., Jang et al. (2023)), our results are independent of any previous findings from regret guarantees of online gambling algorithms. Additionally, our newly derived MI-based bound recovers many previous results and improves our understanding of their potential limitations. Finally, we empirically compare various $f$-information measures for generalization, demonstrating the improvement of our new bounds over the previous bounds.

泛化界信息论理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。