让表格数据模型在准确率和硬件成本间自动找最优平衡。
HAPEns: Hardware-Aware Post-Hoc Ensembling for Tabular Data
- 基于多目标优化,在准确率与资源消耗间构建帕累托前沿的集成方法。
- 83个数据集实验表明,显著优于基线,尤其在内存使用上表现突出。
- 适合追求部署效率的工业级应用,也适用于资源受限场景。
集成学习广泛用于表格数据以提升预测性能与鲁棒性,但更大集成往往带来更高的硬件需求。我们提出HAPEns,一种后处理集成方法,显式平衡准确率与硬件效率。受多目标优化和质量多样性启发,HAPEns沿预测性能与资源消耗的帕累托前沿构建多样化集成组合。现有硬件感知后处理集成基线缺失,凸显本方法的新颖性。在83个表格分类数据集上的实验显示,HAPEns显著优于基线,能发现更优的集成性能与部署成本权衡。消融研究进一步表明,内存使用是特别有效的目标度量。此外,即使使用贪心集成算法,通过静态多目标加权也能显著提升性能。
原文摘要 · Abstract (English)
Ensembling is commonly used in machine learning on tabular data to boost predictive performance and robustness, but larger ensembles often lead to increased hardware demand. We introduce HAPEns, a post-hoc ensembling method that explicitly balances accuracy against hardware efficiency. Inspired by multi-objective and quality diversity optimization, HAPEns constructs a diverse set of ensembles along the Pareto front of predictive performance and resource usage. Existing hardware-aware post-hoc ensembling baselines are not available, highlighting the novelty of our approach. Experiments on 83 tabular classification datasets show that HAPEns significantly outperforms baselines, finding superior trade-offs for ensemble performance and deployment cost. Ablation studies also reveal that memory usage is a particularly effective objective metric. Further, we show that even a greedy ensembling algorithm can be significantly improved in this task with static multi-objective weighting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。