arXiv:2504.13218cs.LGcs.AI2025-04被引 4

提出统一框架,让模型连续学习不同模态数据而不遗忘旧知识。

Harmony: A Unified Framework for Modality Incremental Learning

  • 设计自适应特征调制与模态桥梁机制,实现跨模态对齐。
  • 在仅单模态数据条件下,仍能保持知识连续性并准确完成多模态任务。
  • 适用于真实世界中不断出现新模态的持续学习场景。

增量学习旨在使模型能从持续演变的数据流中不断获取知识,同时保留已有能力。现有研究主要关注单模态或固定模态的多模态增量学习,但现实场景中常出现全新模态,带来额外挑战。本文探讨构建统一模型以实现连续模态序列的增量学习可行性,提出一种新范式——模态增量学习(MIL),每个阶段涉及不同模态的数据。为此,我们提出名为Harmony的新框架,通过自适应兼容特征调制和累积模态桥接机制,构建历史模态特征,实现模态知识积累与对齐,有效缓解模态差异,支持在统一框架下完成多模态任务。即使每个阶段仅有单模态数据,该方法仍可保持知识留存。大量实验表明,该方法显著优于现有增量学习方法,在MIL任务中验证了其有效性。

原文摘要 · Abstract (English)

Incremental learning aims to enable models to continuously acquire knowledge from evolving data streams while preserving previously learned capabilities. While current research predominantly focuses on unimodal incremental learning and multimodal incremental learning where the modalities are consistent, real-world scenarios often present data from entirely new modalities, posing additional challenges. This paper investigates the feasibility of developing a unified model capable of incremental learning across continuously evolving modal sequences. To this end, we introduce a novel paradigm called Modality Incremental Learning (MIL), where each learning stage involves data from distinct modalities. To address this task, we propose a novel framework named Harmony, designed to achieve modal alignment and knowledge retention, enabling the model to reduce the modal discrepancy and learn from a sequence of distinct modalities, ultimately completing tasks across multiple modalities within a unified framework. Our approach introduces the adaptive compatible feature modulation and cumulative modal bridging. Through constructing historical modal features and performing modal knowledge accumulation and alignment, the proposed components collaboratively bridge modal differences and maintain knowledge retention, even with solely unimodal data available at each learning phase.These components work in concert to establish effective modality connections and maintain knowledge retention, even when only unimodal data is available at each learning stage. Extensive experiments on the MIL task demonstrate that our proposed method significantly outperforms existing incremental learning methods, validating its effectiveness in MIL scenarios.

增量学习多模态知识保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。