arXiv:2509.18962cs.LG2025-09被引 1

提出可自适应选模型的绿色在线集成方法,兼顾精度与资源消耗。

Lift What You Can: Green Online Learning with Heterogeneous Ensembles

  • 基于资源约束从多样超参数模型中动态选择子集训练
  • ζ-策略在保持高精度的同时显著降低计算资源使用
  • 适合注重可持续性的实时数据流场景,如边缘计算

面向数据流挖掘的集成方法需持续管理多个模型并随分布变化更新。然而现有方法过度关注预测性能,忽视各模型的计算开销,难以满足可持续性需求。为此,我们提出异构在线集成框架HEROS:每轮训练在初始化时采用不同超参数的模型池中,依据资源限制选择子集进行更新。通过马尔可夫决策过程建模性能与可持续性间的权衡,并设计多种选择策略。尤其提出新型ζ-策略,聚焦以更低开销训练近最优模型。基于随机模型的理论证明显示,该策略在逼近最优性能的同时资源消耗更少。11个基准数据集上的实验表明,ζ-策略在多数情况下性能优于或媲美先进方法,且资源效率显著提升。

原文摘要 · Abstract (English)

Ensemble methods for stream mining necessitate managing multiple models and updating them as data distributions evolve. Considering the calls for more sustainability, established methods are however not sufficiently considerate of ensemble members' computational expenses and instead overly focus on predictive capabilities. To address these challenges and enable green online learning, we propose heterogeneous online ensembles (HEROS). For every training step, HEROS chooses a subset of models from a pool of models initialized with diverse hyperparameter choices under resource constraints to train. We introduce a Markov decision process to theoretically capture the trade-offs between predictive performance and sustainability constraints. Based on this framework, we present different policies for choosing which models to train on incoming data. Most notably, we propose the novel $ζ$-policy, which focuses on training near-optimal models at reduced costs. Using a stochastic model, we theoretically prove that our $ζ$-policy achieves near optimal performance while using fewer resources compared to the best performing policy. In our experiments across 11 benchmark datasets, we find empiric evidence that our $ζ$-policy is a strong contribution to the state-of-the-art, demonstrating highly accurate performance, in some cases even outperforming competitors, and simultaneously being much more resource-friendly.

在线学习集成方法绿色计算资源优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。