arXiv:2410.19925cs.CLcs.CV2024-10被引 5

用持续学习提升多模态大模型,让视觉能力增强同时不丢语言能力。

Improving Multimodal Large Language Models Using Continual Learning

  • 将多模态融合视为持续学习问题,设计新方法防止语言能力退化。
  • 相比原版LLaVA,语言性能下降减少15%,多模态准确率保持高位。
  • 适合需要持续学习新视觉任务又不牺牲语言能力的研究者。

生成式大语言模型(LLM)具备强大能力,通过整合预训练视觉模型可构建多模态大语言模型(MLLM)。然而,这种融合常导致自然语言理解与生成任务性能显著下降。本研究以LLaVA MLLM为例,将此问题视为持续学习挑战,评估五种持续学习方法以缓解遗忘。结果发现一种新方法能有效提升视觉理解,同时最小化语言性能损失。该方法使语言性能下降比原始LLaVA方案减少最多15%,且维持高多模态准确率。我们进一步在一系列视觉-语言任务上验证其鲁棒性,成功在保留语言技能的同时获得新的多模态能力。

原文摘要 · Abstract (English)

Generative large language models (LLMs) exhibit impressive capabilities, which can be further augmented by integrating a pre-trained vision model into the original LLM to create a multimodal LLM (MLLM). However, this integration often significantly decreases performance on natural language understanding and generation tasks, compared to the original LLM. This study investigates this issue using the LLaVA MLLM, treating the integration as a continual learning problem. We evaluate five continual learning methods to mitigate forgetting and identify a technique that enhances visual understanding while minimizing linguistic performance loss. Our approach reduces linguistic performance degradation by up to 15% over the LLaVA recipe, while maintaining high multimodal accuracy. We also demonstrate the robustness of our method through continual learning on a sequence of vision-language tasks, effectively preserving linguistic skills while acquiring new multimodal capabilities. Project webpage: https://shikhar-srivastava.github.io/cl-for-improving-mllms

多模态持续学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。