对比56个表格分类任务,发现混合投票法在多分类中表现最佳。
Choosing a parallel heterogeneous ensemble method for tabular classification

- 采用并行集成方法,结合软投票与硬投票策略提升分类性能。
- 在56个任务上验证,混合投票法优于单一模型和传统集成方法。
- 适合需要高精度多分类的表格数据应用场景。
在来自OpenML CC18的56个中小型表格分类任务上对比了并行集成方法,并据此得出一套“最佳实践”建议。该建议在28个额外任务上通过TabArena预计算数据验证,显著优于单个最优模型,且达到或超过个别集成方法的表现。关键发现:一是融合(Blending)与堆叠(Stacking)的不一致性相互独立,发生在不同任务上;二是尽管硬投票的概率分类表现较弱(因使用投票比例作为后验估计),但鲁棒软投票(Robust Soft Voting)在多分类情形下尤为成功。
原文摘要 · Abstract (English)
Parallel ensemble methods were compared on $56$ small-to-medium tabular classification tasks drawn from OpenML CC18. A set of ``best practice'' recommendations on the use of ensemble methods was derived from these observations. It was later validated on 28 additional tasks using TabArena's precomputed data, where the recommendation set significantly outperformed Single Best and matched or exceeded individual ensemble methods. Two key observations were made. First, Blending and Stacking are inconsistent, but their inconsistencies are independent and happen on different tasks. Second, while Hard Voting's probabilistic classification is rather weak, a consequence of using vote proportions as posterior estimates, Robust Soft Voting's probabilistic classification is particularly successful, especially in the multiclass case.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。