arXiv:2602.06773cs.LGstat.ML2026-02

提出多校准梯度提升的收敛性理论,解释其为何在实际中表现良好。

On the Convergence of Multicalibration Gradient Boosting

  • 通过分析预测更新幅度衰减速度,建立收敛性保证。
  • 证明经验多校准误差以 $O(1/\sqrt{T})$ 速率下降,平滑条件下可线性收敛。
  • 实验验证理论,揭示快速收敛的实际适用场景。

多校准梯度提升最近成为一种可扩展的方法,能实证生成近似多校准的预测器,并已在网页规模部署。尽管有此实证成功,其收敛性质尚不明确。本文为多校准梯度提升算法提供了计算保证:我们证明连续预测更新的幅度以 $O(1/\\/sqrt{T})$ 速度衰减,这意味着经验多校准误差在轮次上具有相同的收敛速率界。在弱学习器满足额外光滑性假设下,该速率可提升至线性收敛。我们进一步建立了自适应变体的收敛性。在真实数据集上的实验支持了理论,并阐明了该方法实现快速收敛的条件。

原文摘要 · Abstract (English)

Multicalibration gradient boosting has recently emerged as a scalable method that empirically produces approximately multicalibrated predictors and has been deployed at web scale. Despite this empirical success, its convergence properties are not well understood. In this paper, we provide computational guarantees for multicalibration gradient boosting algorithms. We show that the magnitude of successive prediction updates decays at $O(1/\sqrt{T})$, which implies the same convergence rate bound for the empirical multicalibration error over rounds. Under additional smoothness assumptions on the weak learners, this rate improves to linear convergence. We further establish convergence for adaptive variants. Experiments on real-world datasets support our theory and clarify the regimes in which the method achieves fast convergence.

梯度提升多校准收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。