用标签翻转检测流数据概念漂移,发现决策树集成效果差。
Pitfalls of Unlabeled Disagreement-Based Drift Detection in Streaming Tree Ensembles

- 通过翻转标签构造集成模型间分歧度量
- 在决策树集成上漂移检测准确率低于基于损失的检测方法
- 适合研究流数据中模型可塑性与检测性能关系的人
在高速数据流中检测概念漂移仍具挑战性,尤其当模型需在无标签数据上运行且避免由良性变化引发的误报时。尽管基于分歧的不确定性在神经网络中表现良好,但其在增量决策树集成(IDTs)中的应用尚未深入探索。本文通过在集成成员中进行标签翻转,构建批次特定的分歧度量,并评估其在表格数据流中的漂移检测效果。实验表明,该方法在多层感知机(MLPs)集成中表现良好,但在IDTs上始终劣于基于损失的检测器。我们归因于IDTs的内在刚性:主要通过结构扩展学习,参数调整有限,导致模型可塑性不足,分歧无法可靠反映学习潜力。近期利用非重叠规则分解重构IDTs的研究为提升适应性提供了新方向。
原文摘要 · Abstract (English)
Detecting concept drift in high-speed data streams remains challenging, particularly when models must operate on unlabeled data and avoid false alarms caused by benign shifts. While disagreement-based uncertainty has shown promise in neural networks, its adaptation to ensembles of incremental decision trees (IDTs) remains largely unexplored. We investigate this approach by constructing batch-specific disagreement measures via label flipping in ensemble members and evaluating their effectiveness for drift detection in tabular data streams. Our experiments show that, although this method performs well in ensembles of multi-layer perceptrons (MLPs), it consistently underperforms loss-based detectors when applied to IDTs. We attribute this behavior to the intrinsic rigidity of IDTs: learning primarily through structural expansion, with limited parameter adaptation, restricts model plasticity and prevents disagreement from reliably reflecting learning potential. Recent work on restructuring IDTs using their intrinsic decomposition into non-overlapping rules offers a promising direction for improving adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。