arXiv:2412.09035cs.LG2024-12被引 1

用遗传算法动态优化模型组合,应对数据概念漂移问题。

Pulling the Carpet Below the Learner's Feet: Genetic Algorithm To Learn Ensemble Machine Learning Model During Concept Drift

  • 设计两级集成模型,结合全局模型与漂移检测器
  • 在未知漂移场景下性能优于单模型方法
  • 适合需要自适应更新的实时机器学习系统

近年来,数据驱动模型尤其是机器学习(ML)模型在科学与工程领域广泛应用。但在真实动态环境中使用时,用户常面临概念漂移(CD)挑战。本文探索遗传算法(GA)在应对此类问题中的应用,提出一种新型两级集成机器学习模型:该模型包含一个全局模型与一个概念漂移检测器,作为种群中多个机器学习流水线模型的聚合器;每个子模型自带可调漂移检测器,负责自主重训练其本地模型。此外,我们还证明可通过引入现成的自动机器学习方法进一步提升性能。通过大量合成数据集分析,结果表明该模型在未知漂移特征的场景下显著优于单一机器学习流水线搭配漂移处理机制。总体而言,本研究展示了通过启发式与自适应优化过程(如遗传算法)构建的集成学习与漂移检测模型,在处理复杂漂移事件方面的潜力。

原文摘要 · Abstract (English)

Data-driven models, in general, and machine learning (ML) models, in particular, have gained popularity over recent years with an increased usage of such models across the scientific and engineering domains. When using ML models in realistic and dynamic environments, users need to often handle the challenge of concept drift (CD). In this study, we explore the application of genetic algorithms (GAs) to address the challenges posed by CD in such settings. We propose a novel two-level ensemble ML model, which combines a global ML model with a CD detector, operating as an aggregator for a population of ML pipeline models, each one with an adjusted CD detector by itself responsible for re-training its ML model. In addition, we show one can further improve the proposed model by utilizing off-the-shelf automatic ML methods. Through extensive synthetic dataset analysis, we show that the proposed model outperforms a single ML pipeline with a CD algorithm, particularly in scenarios with unknown CD characteristics. Overall, this study highlights the potential of ensemble ML and CD models obtained through a heuristic and adaptive optimization process such as the GA one to handle complex CD events.

概念漂移集成学习遗传算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。