arXiv:2503.07663cs.LGcs.AI2025-03EMNLP被引 6

解决多模态大模型增量学习中的遗忘与对齐问题。

Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMs

  • 先合并后重对齐,简单有效提升多模态模型增量学习性能。
  • 扩展至四种模态时,平均后向相对收益达99.84%,近乎无损。
  • 无需改架构、不增参数,适合快速部署于现有大模型系统。

多模态大语言模型(MLLMs)在集成越来越多模态后展现出更强的通用性。由于训练成本高昂,通过模态增量持续学习(MCL)复用已有模型并扩展新模态更具效率。当前MCL研究尚处于起步阶段。本文深入分析了MCL中性能下降的原因,发现其不仅存在传统持续学习中的遗忘问题,还存在模态无关与模态特定组件间的错位问题。为此,提出一种简洁高效的MCL范式——“合并后重对齐”(MERA),同时缓解遗忘与错位。MERA不引入额外模型开销,也不修改模型结构,易于部署且高度可复用。大量实验表明,MERA在扩展至四种模态时,平均后向相对收益达99.84%,实现近乎无损的持续学习性能。研究揭示了MCL中的对齐问题,并展示了如何在持续学习中调节MLLM不同组件。

原文摘要 · Abstract (English)

Recent advances in Multimodal Large Language Models (MLLMs) have enhanced their versatility as they integrate a growing number of modalities. Considering the heavy cost of training MLLMs, it is efficient to reuse the existing ones and extend them to more modalities through Modality-incremental Continual Learning (MCL). The exploration of MCL is in its early stages. In this work, we dive into the causes of performance degradation in MCL. We uncover that it suffers not only from forgetting as in traditional continual learning, but also from misalignment between the modality-agnostic and modality-specific components. To this end, we propose an elegantly simple MCL paradigm called "MErge then ReAlign" (MERA) to address both forgetting and misalignment. MERA avoids introducing heavy model budgets or modifying model architectures, hence is easy to deploy and highly reusable in the MLLM community. Extensive experiments demonstrate the impressive performance of MERA, holding an average of 99.84\% Backward Relative Gain when extending to four modalities, achieving nearly lossless MCL performance. Our findings underscore the misalignment issue in MCL. More broadly, our work showcases how to adjust different components of MLLMs during continual learning.

多模态持续学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。