arXiv:2602.18055cs.LG2026-02

提出统一多模态持续学习框架,解决模型学新知识时遗忘旧知识的问题。

Continual-NExT: A Unified Comprehension And Generation Continual Learning Framework

  • 设计MAGE方法,融合通用与专家LoRA提升跨模态知识迁移
  • 在多个任务上实现最佳持续学习表现,显著减少遗忘和幻觉
  • 适合需要长期适应新场景的多模态大模型研究者使用

双模态多模态大语言模型(Dual-to-Dual MLLMs)可通过文本与图像实现统一的多模态理解与生成。尽管具备强大的即时学习与泛化能力,但在持续进化方面仍存在不足,严重影响其对动态现实场景的适应性。主要挑战包括:学习新任务不可避免地破坏已有知识,除传统灾难性遗忘外,还面临幻觉、指令偏离及跨模态知识迁移失败等问题。目前尚未建立针对此类模型的标准化持续学习框架,导致这些问题未被系统研究。为此,本文提出Continual-NExT框架,专为双模态多模态大模型设计,并配备精心构建的评估指标。为提升持续学习能力,我们提出高效MAGE(General LoRA与Expert LoRA的混合与聚合)方法,进一步促进跨模态知识迁移并缓解遗忘。大量实验表明,MAGE优于其他持续学习方法,在多项任务上达到当前最优性能。

原文摘要 · Abstract (English)

Dual-to-Dual MLLMs refer to Multimodal Large Language Models, which can enable unified multimodal comprehension and generation through text and image modalities. Although exhibiting strong instantaneous learning and generalization capabilities, Dual-to-Dual MLLMs still remain deficient in lifelong evolution, significantly affecting continual adaptation to dynamic real-world scenarios. One of the challenges is that learning new tasks inevitably destroys the learned knowledge. Beyond traditional catastrophic forgetting, Dual-to-Dual MLLMs face other challenges, including hallucination, instruction unfollowing, and failures in cross-modal knowledge transfer. However, no standardized continual learning framework for Dual-to-Dual MLLMs has been established yet, leaving these challenges unexplored. Thus, in this paper, we establish Continual-NExT, a continual learning framework for Dual-to-Dual MLLMs with deliberately-architected evaluation metrics. To improve the continual learning capability of Dual-to-Dual MLLMs, we propose an efficient MAGE (Mixture and Aggregation of General LoRA and Expert LoRA) method to further facilitate knowledge transfer across modalities and mitigate forgetting. Extensive experiments demonstrate that MAGE outperforms other continual learning methods and achieves state-of-the-art performance.

持续学习多模态大模型知识迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。