多模态持续学习新框架,有效防遗忘并融合跨模态信息
Multi-Modal Continual Learning via Cross-Modality Adapters and Representation Alignment with Knowledge Preservation
- 用混合专家结构的跨模态适配器融合多源输入
- 引入表示对齐损失,提升多模态表征鲁棒性
- 适合需要长期学习多模态数据的场景
持续学习使模型能适应新任务同时保留旧知识。现有方法多集中于单模态数据,而多模态学习利用多样化感官输入,更接近人类感知。然而,多模态持续学习面临新信息整合与灾难性遗忘双重挑战。本文提出基于预训练模型的多模态持续学习框架,包含新颖的跨模态适配器(采用混合专家结构)以促进跨任务的多模态信息融合;引入表示对齐损失,增强多模态表征的鲁棒性,并通过正则化学习到的表示关系来保护历史知识。在多个多模态数据集上的实验表明,该方法在类别增量与领域增量学习中均显著优于基线,准确率更高且遗忘更少。
原文摘要 · Abstract (English)
Continual learning is essential for adapting models to new tasks while retaining previously acquired knowledge. While existing approaches predominantly focus on uni-modal data, multi-modal learning offers substantial benefits by utilizing diverse sensory inputs, akin to human perception. However, multi-modal continual learning presents additional challenges, as the model must effectively integrate new information from various modalities while preventing catastrophic forgetting. In this work, we propose a pre-trained model-based framework for multi-modal continual learning. Our framework includes a novel cross-modality adapter with a mixture-of-experts structure to facilitate effective integration of multi-modal information across tasks. We also introduce a representation alignment loss that fosters learning of robust multi-modal representations, and regularize relationships between learned representations to preserve knowledge from previous tasks. Experiments on several multi-modal datasets demonstrate that our approach consistently outperforms baselines in both class-incremental and domain-incremental learning, achieving higher accuracy and reduced forgetting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。