arXiv:2510.22405cs.LGcs.AI2025-10中稿 · IEEE CogMI 2025 - …被引 1

用外部知识增强记忆缓冲,解决在线行为分析模型随时间退化问题

Knowledge-guided Continual Learning for Behavioral Analytics Systems

  • 用外部知识生成新数据,动态扩充固定大小的旧数据缓冲区
  • 在三个偏移行为分类数据集上,性能优于传统重放方法
  • 适合长期运行的在线内容监测系统开发者使用

在线平台上的用户行为持续演化,反映现实世界中信息发布方式的变化,如有益内容或仇恨言论的增减。依赖此类内容训练的模型可能因数据漂移导致性能下降,进而使行为分析系统失效。然而,直接通过新数据微调模型易引发灾难性遗忘。基于重放的持续学习方法通过保留过往任务的重要训练样本,可有效缓解遗忘问题,但其核心局限在于缓冲区容量固定。本文提出一种基于外部知识的数据增强策略,嵌入重放式持续学习框架,以突破缓冲区容量限制。我们在三个先前研究中用于异常行为分类的数据集上评估了多种策略,结果表明,引入外部知识进行数据增强能显著提升模型性能,优于基准重放方法。

原文摘要 · Abstract (English)

User behavior on online platforms is evolving, reflecting real-world changes in how people post, whether it's helpful messages or hate speech. Models that learn to capture this content can experience a decrease in performance over time due to data drift, which can lead to ineffective behavioral analytics systems. However, fine-tuning such a model over time with new data can be detrimental due to catastrophic forgetting. Replay-based approaches in continual learning offer a simple yet efficient method to update such models, minimizing forgetting by maintaining a buffer of important training instances from past learned tasks. However, the main limitation of this approach is the fixed size of the buffer. External knowledge bases can be utilized to overcome this limitation through data augmentation. We propose a novel augmentation-based approach to incorporate external knowledge in the replay-based continual learning framework. We evaluate several strategies with three datasets from prior studies related to deviant behavior classification to assess the integration of external knowledge in continual learning and demonstrate that augmentation helps outperform baseline replay-based approaches.

持续学习行为分析知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。