arXiv:2507.18242cs.LG2025-07被引 1

对比六种线性规划提升法,发现其可生成更稀疏且性能不俗的集成模型。

Boosting Revisited: Benchmarking and Advancing LP-Based Ensemble Methods

  • 用线性规划重构提升算法,实现完全校正的集成学习。
  • 在浅树条件下超越或媲美XGBoost、LightGBM,且模型更稀疏。
  • 适合追求模型简洁性与可解释性的机器学习研究者。

尽管基于线性规划的完全校正提升方法具有理论优势,但其在实际中的应用仍有限。本文首次在20个多样化数据集上大规模评估六种基于线性规划的提升方法,包括两种新提出的方法(NM-Boost 和 QRLP-Boost)。我们考察了启发式与最优基学习器的使用效果,并分析了准确率、集成稀疏性、间隔分布、即时性能及超参数敏感性。结果表明,当使用浅层决策树时,完全校正方法可超越或媲美当前主流启发式方法(如XGBoost、LightGBM),同时生成显著更稀疏的集成模型。此外,这些方法还能在不损失性能的前提下对预训练集成模型进行剪枝。研究还揭示了在该框架中使用最优决策树的优势与局限。

原文摘要 · Abstract (English)

Despite their theoretical appeal, totally corrective boosting methods based on linear programming have received limited empirical attention. In this paper, we conduct the first large-scale experimental study of six LP-based boosting formulations, including two novel methods, NM-Boost and QRLP-Boost, across 20 diverse datasets. We evaluate the use of both heuristic and optimal base learners within these formulations, and analyze not only accuracy, but also ensemble sparsity, margin distribution, anytime performance, and hyperparameter sensitivity. We show that totally corrective methods can outperform or match state-of-the-art heuristics like XGBoost and LightGBM when using shallow trees, while producing significantly sparser ensembles. We further show that these methods can thin pre-trained ensembles without sacrificing performance, and we highlight both the strengths and limitations of using optimal decision trees in this context.

集成学习线性规划模型稀疏性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。