提出窗口困境:漂移检测结果受窗口选择影响,未必反映真实数据变化。
The Window Dilemma: Why Concept Drift Detection is Ill-Posed
- 通过对比不同窗口期数据判断漂移,但窗口选择本身会误导结果
- 实验证明传统批量学习常优于漂移感知方法
- 质疑漂移检测在实际中的可验证性与必要性
数据流中数据生成过程的非平稳性导致分布随时间变化,这一现象称为概念漂移,已被广泛研究。当前漂移检测器主要通过比较数据流中不同窗口区域的差异来识别漂移。本文提出“窗口困境”:感知到的漂移可能源于窗口设置,而非真实数据生成过程的变化。同时指出漂移检测本质上是病态问题,因实际中难以验证漂移事件的真实性。通过示例和多种自适应策略的实证比较,发现传统批量学习方法往往优于漂移感知模型,从而引发对漂移检测器在流分类中作用的质疑。
原文摘要 · Abstract (English)
Non-stationarity of an underlying data generating process that leads to distributional changes over time is a key characteristic of Data Streams. This phenomenon, commonly referred to as Concept Drift, has been intensively studied, and Concept Drift Detectors have been established as a class of methods for detecting such changes (drifts). For the most part, Drift Detectors compare regions (windows) of the data stream and detect drift if those windows are sufficiently dissimilar. In this work, we introduce the Window Dilemma, an observation that perceived drift is a product of windowing and not necessarily the underlying data generating process. Additionally, we highlight that drift detection is ill-posed, primarily because verification of drift events are implausible in practice. We demonstrate these contributions first by an illustrative example, followed by empirical comparisons of drift detectors against a variety of alternative adaptation strategies. Our main finding is that traditional batch learning techniques often perform better than their drift-aware counterparts further bringing into question the purpose of detectors in Stream Classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。