arXiv:2606.22026cs.LGcs.AI2026-06

提出新方法模拟数据分布漂移,高效适配模型变化。

Cluster-Specific Localized Drift Detection for Efficient Batch Model Adaptation under Controlled Distribution Shift

论文配图:Cluster-Specific Localized Drift Detection for Efficient Batch Model Adaptation under Controlled Distribution Shift
图 1 · 摘自论文原文
  • 通过特征空间分块施加结构化扰动,生成可控数据流
  • 在5个数据集上验证6种策略,发现局部漂移检测更优
  • 适合需持续更新的在线学习系统,如金融风控

部署于动态环境中的机器学习系统常面临非平稳数据分布,可控的分布漂移会逐步降低预测性能。然而,多数常用表格基准数据集缺乏明确的时间结构,限制了漂移适应方法的可复现评估。本文提出一种基于聚类的分布漂移模拟框架,通过在特征空间分区上施加结构化扰动,将静态表格数据转换为可控的演化数据流。基于该框架,系统评估了六种适应策略:静态学习、滑动窗口重训练、全局ADWIN重训练、聚类局部ADWIN重训练、随机子空间漂移检测和特征分区漂移检测。实验在五个涵盖分类与回归任务的基准数据集上进行,使用线性模型、k-近邻、树集成、提升方法及自适应在线学习器等多样模型族。

原文摘要 · Abstract (English)

Machine learning systems deployed in dynamic environments frequently operate under nonstationary data distributions, where controlled distribution shift can progressively degrade predictive performance. However, many widely used tabular benchmark datasets lack explicit temporal structure, limiting reproducible evaluation of drift adaptation methods. This work proposes a cluster-induced distribution shift simulation framework that transforms static tabular datasets into controlled evolving data streams through structured perturbations across featurespace partitions. Using this framework, six adaptation strategies are systematically evaluated: static learning, sliding-window retraining, global ADWIN retraining, cluster-local ADWIN retraining, random subspace drift detection, and feature-partitioned drift detection. Experiments are conducted on five benchmark datasets covering both classification and regression tasks using diverse predictive model families, including linear models, k-Nearest Neighbours, tree ensembles, boosting methods, and adaptive online learners.

分布漂移在线学习模型适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。