针对复杂不平衡数据流,改进在线袋装算法以提升少数类分类性能。
Improving Online Bagging for Complex Imbalanced Data Stream
- 引入邻域欠采样与过采样机制,识别并处理少数类中的难例
- 在合成数据流上表现优于传统在线袋装方法,准确率显著提升
- 适合处理概念漂移和少数类细粒度分解的复杂场景
从不平衡且存在概念漂移的数据流中学习分类器仍具挑战性。现有方法多仅关注全局不平衡比率的变化,忽略局部难度因素,如少数类细分为子概念、边界或稀有样本的存在。这些因素会降低主流在线分类器性能。本文提出改进的重采样在线袋装方法,即邻域欠采样/过采样在线袋装(Neighbourhood Undersampling/Oversampling Online Bagging),更有效应对少数类中的不安全样本。在合成复杂不平衡数据流上的计算实验表明,该方法优于早期在线袋装重采样集成方法。
原文摘要 · Abstract (English)
Learning classifiers from imbalanced and concept drifting data streams is still a challenge. Most of the current proposals focus on taking into account changes in the global imbalance ratio only and ignore the local difficulty factors, such as the minority class decomposition into sub-concepts and the presence of unsafe types of examples (borderline or rare ones). As the above factors present in the stream may deteriorate the performance of popular online classifiers, we propose extensions of resampling online bagging, namely Neighbourhood Undersampling or Oversampling Online Bagging to take better account of the presence of unsafe minority examples. The performed computational experiments with synthetic complex imbalanced data streams have shown their advantage over earlier variants of online bagging resampling ensembles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。