解决多模态大模型在跨模态任务中持续学习时的遗忘问题。
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
- 通过伪目标生成缓解任务类型切换导致的遗忘。
- 引入指令知识蒸馏,保留旧模态处理能力。
- 在6个跨模态任务上显著优于现有方法。
本文提出多模态不一致持续学习(MICL),一种针对多模态大语言模型(MLLMs)的新持续学习场景,涉及图像、音频或视频等不一致模态及标注或问答等不同任务类型。与现有的仅视觉或模态增量设置不同,MICL同时包含模态和任务类型变化,二者均引发灾难性遗忘。为此,我们提出MoInCL,采用伪目标生成模块减轻已见模态中任务类型切换带来的遗忘;同时引入基于指令的知识蒸馏,保留新模态引入时对旧模态的处理能力。我们在总计六个任务上评估MICL,实验验证了MoInCL的有效性,结果表明其显著优于代表性及前沿持续学习基线。
原文摘要 · Abstract (English)
In this paper, we introduce Modality-Inconsistent Continual Learning (MICL), a new continual learning scenario for Multimodal Large Language Models (MLLMs) that involves tasks with inconsistent modalities (image, audio, or video) and varying task types (captioning or question-answering). Unlike existing vision-only or modality-incremental settings, MICL combines modality and task type shifts, both of which drive catastrophic forgetting. To address these challenges, we propose MoInCL, which employs a Pseudo Targets Generation Module to mitigate forgetting caused by task type shifts in previously seen modalities. It also incorporates Instruction-based Knowledge Distillation to preserve the model's ability to handle previously learned modalities when new ones are introduced. We benchmark MICL using a total of six tasks and conduct experiments to validate the effectiveness of our MoInCL. The experimental results highlight the superiority of MoInCL, showing significant improvements over representative and state-of-the-art continual learning baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。