arXiv:2607.26838cs.LG2026-07

给出AdaBoost泛化误差的紧致上界,揭示其与弱学习器优势和样本量的关系。

Tight Generalization Bound for AdaBoost

  • 基于投票函数的0-经验γ/2边缘损失,结合新提出的边缘泛化界
  • 泛化误差上界为Θ((d ln(nγ²/d))/(nγ²) + ln(1/δ)/n)
  • 对理论分析者有参考价值,适合关注提升学习泛化性的研究者

本文证明了AdaBoost的泛化误差为Θ((d ln(nγ²/d))/(nγ²) + ln(1/δ)/n),其中γ是弱学习器保证的优势,d是包含弱假设的类的VC维,n是样本数量,δ是置信参数。本文贡献在于给出了该上界;匹配的下界已有前人工作证明。上界证明基于已知事实:AdaBoost输出的投票分类器在经验γ/2边缘损失上为零,并结合了我们认为是首个针对投票分类器的基于边缘的泛化界。

原文摘要 · Abstract (English)

In this paper we show that the generalization error of AdaBoost is $Θ\big(\tfrac{d\ln(nγ^{2}/d)}{nγ^2}+\tfrac{\ln(1/δ)}{n}\big)$, where $γ$ is the advantage guaranteed by the weak learner, $d$ is the VC-dimension of the class containing the weak hypotheses, $n$ is the sample size, and $δ$ is the confidence parameter. The contribution of this paper is the upper bound; the matching lower bound follows from prior work. The upper bound proof follows by combining the known fact that AdaBoost outputs a voting classifier whose voting function has zero empirical $γ/2$-margin loss with what is, to the best of our knowledge, a new margin-based generalization bound for voting classifiers.

AdaBoost泛化边界统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。