arXiv:2504.04613cs.LG2025-04被引 5

用双增强孪生网络实现数据流中低预算主动学习,提升实时学习效率。

SiameseDuo++: Active Learning from Data Streams with Dual Augmented Siamese Networks

  • 双孪生网络协同学习,基于潜在空间的主动采样与数据增强
  • 在有限标注预算下,相比基线方法提升学习速度与性能
  • 适合需要实时处理高维数据流的场景,如监控与推荐系统

数据流挖掘(stream learning)是处理高速到达数据的学习领域,广泛应用于关键基础设施监控、社交媒体分析和推荐系统。其研究挑战包括数据分布随时间变化(概念漂移)、数据流通常缺乏真实标签,以及需在有限内存下实时处理大量数据。本文提出SiameseDuo++方法,通过主动学习在有限标注预算内自动选择待标注样本。该方法增量训练两个协同工作的孪生神经网络,并在潜在空间中进行数据增强。主动学习策略与数据增强均在潜在空间中实现。实验表明,该方法在学习速度和/或性能上优于多个强基线及前沿方法。为促进开放科学,代码与数据集已公开。

原文摘要 · Abstract (English)

Data stream mining, also known as stream learning, is a growing area which deals with learning from high-speed arriving data. Its relevance has surged recently due to its wide range of applicability, such as, critical infrastructure monitoring, social media analysis, and recommender systems. The design of stream learning methods faces significant research challenges; from the nonstationary nature of the data (referred to as concept drift) and the fact that data streams are typically not annotated with the ground truth, to the requirement that such methods should process large amounts of data in real-time with limited memory. This work proposes the SiameseDuo++ method, which uses active learning to automatically select instances for a human expert to label according to a budget. Specifically, it incrementally trains two siamese neural networks which operate in synergy, augmented by generated examples. Both the proposed active learning strategy and augmentation operate in the latent space. SiameseDuo++ addresses the aforementioned challenges by operating with limited memory and limited labelling budget. Simulation experiments show that the proposed method outperforms strong baselines and state-of-the-art methods in terms of learning speed and/or performance. To promote open science we publicly release our code and datasets.

数据流主动学习孪生网络实时学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。