arXiv:2501.08878cs.LGcs.AI2025-01被引 1

让模型持续学习多个领域数据,避免遗忘且提升泛化能力。

Incrementally Learning Multiple Diverse Data Domains via Multi-Source Dynamic Expansion Model

  • 用多源动态扩展模型,逐步构建新专家应对新任务。
  • 动态注意力机制加速新任务学习,跨领域知识利用率更高。
  • 适合需要长期适应多领域数据的智能系统应用。

持续学习旨在让模型在不断接收新信息的同时保留已有知识。然而,现有研究大多局限于单一数据域的简单学习场景。本文转向更复杂、更真实的多源数据域学习环境,提出多源动态扩展模型(MSDEM),利用预训练模型作为骨干网络,逐步建立新专家以适应新任务。同时设计动态可扩展注意力机制,有选择性地融合多骨干知识,加速新任务学习;引入动态图权重路由器,策略性重用先前所有参数与表示,最大化正向知识迁移,进一步提升泛化性能。大量实验表明,所提方法达到当前最优效果。

原文摘要 · Abstract (English)

Continual Learning seeks to develop a model capable of incrementally assimilating new information while retaining prior knowledge. However, current research predominantly addresses a straightforward learning context, wherein all data samples originate from a singular data domain. This paper shifts focus to a more complex and realistic learning environment, characterized by data samples sourced from multiple distinct domains. We tackle this intricate learning challenge by introducing a novel methodology, termed the Multi-Source Dynamic Expansion Model (MSDEM), which leverages various pre-trained models as backbones and progressively establishes new experts based on them to adapt to emerging tasks. Additionally, we propose an innovative dynamic expandable attention mechanism designed to selectively harness knowledge from multiple backbones, thereby accelerating the new task learning. Moreover, we introduce a dynamic graph weight router that strategically reuses all previously acquired parameters and representations for new task learning, maximizing the positive knowledge transfer effect, which further improves generalization performance. We conduct a comprehensive series of experiments, and the empirical findings indicate that our proposed approach achieves state-of-the-art performance.

持续学习多领域动态扩展知识迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。