系统梳理多标签数据流分类的最新方法与挑战
A Systematic Literature Review on Multi-label Data Stream Classification
- 构建了多标签数据流分类方法的完整分类体系
- 分析了现有方法对概念漂移和新标签出现的应对策略
- 适合关注实时多标签学习的研究者与工程师
多标签数据流分类因其在现实场景中的广泛应用而受到广泛关注,但其面临动态环境带来的多重挑战:高速高量的数据持续到达、数据分布变化(概念漂移)、新标签出现(概念演化)以及真实标签到达延迟。本文对多标签数据流分类相关研究进行了系统性综述,深入分析了文献中最新方法的特征,构建了全面的分类层级,探讨了各类方法如何应对上述问题。同时,本文讨论了常用的评估策略,分析了方法的渐近复杂度与资源消耗。最后,识别出当前研究的主要空白,并为未来研究方向提出建议。
原文摘要 · Abstract (English)
Classification in the context of multi-label data streams represents a challenge that has attracted significant attention due to its high real-world applicability. However, this task faces problems inherent to dynamic environments, such as the continuous arrival of data at high speed and volume, changes in the data distribution (concept drift), the emergence of new labels (concept evolution), and the latency in the arrival of ground truth labels. This systematic literature review presents an in-depth analysis of multi-label data stream classification proposals. We characterize the latest methods in the literature, providing a comprehensive overview, building a thorough hierarchy, and discussing how the proposals approach each problem. Furthermore, we discuss the adopted evaluation strategies and analyze the methods' asymptotic complexity and resource consumption. Finally, we identify the main gaps and offer recommendations for future research directions in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。