arXiv:2605.12190stat.MLcs.LG2026-05

为在线学习等序列决策问题提供了新的泛化分析工具。

Information-Theoretic Generalization Bounds for Sequential Decision Making

  • 提出序列超样本框架,分离学习者滤波与证明用的虚拟坐标
  • 泛化误差受序列条件互信息控制,可实现更快收敛速率
  • 适用于在线学习、主动学习和随机多臂赌博机

基于超样本构造的信息论泛化界是批量独立同分布设置下算法依赖泛化分析的核心工具。然而,现有超样本条件互信息(CMI)界无法直接应用于在线学习、流式主动学习和赌博机等序列决策问题,因为数据自适应揭示且学习者沿因果轨迹演化。为此,我们提出一种序列超样本框架,将学习者滤波与证明侧的虚拟坐标扩展相分离。在行交换性假设下,序列泛化差距由序列CMI控制,即每轮选择器-损失信息项之和。我们还建立了伯恩斯坦型改进,在合适方差条件下可获得更快收敛率。选择器-SCMI证明策略适用于在线学习、带重要性加权的流式主动学习以及随机多臂赌博机。

原文摘要 · Abstract (English)

Information-theoretic generalization bounds based on the supersample construction are a central tool for algorithm-dependent generalization analysis in the batch i.i.d.~setting. However, existing supersample conditional mutual information (CMI) bounds do not directly apply to sequential decision-making problems such as online learning, streaming active learning, and bandits, where data are revealed adaptively and the learner evolves along a causal trajectory. To address this limitation, we develop a sequential supersample framework that separates the learner filtration from a proof-side enlargement used for ghost-coordinate comparisons. Under a row-wise exchangeability assumption, the sequential generalization gap is controlled by sequential CMI, a sum of roundwise selector--loss information terms. We also establish a Bernstein-type refinement that yields faster rates under suitable variance conditions. The selector-SCMI proof strategy applies to online learning, streaming active learning with importance weighting, and stochastic multi-armed bandits.

泛化界在线学习信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。