首篇综述多模态大模型持续学习,梳理440篇论文挑战与方向
When Continue Learning Meets Multimodal Large Language Model: A Survey
- 系统分类440篇论文,构建多模态大模型持续学习研究框架
- 揭示模型在新任务中易遗忘旧知识的性能退化问题
- 适合关注大模型动态适应能力的研究者参考
近年来人工智能发展催生了多模态大语言模型(MLLMs),但如何高效适应动态数据分布和多样任务仍面临挑战。针对特定任务微调常导致模型原有知识域性能下降,即‘灾难性遗忘’问题。尽管该问题在持续学习(CL)领域已有广泛研究,但在多模态大模型中呈现新挑战。本文作为首个关于MLLM持续学习的综述,系统分析了440篇相关研究。文章分为四部分:首先综述MLLM最新进展,涵盖模型创新、基准测试及多领域应用;其次分类梳理持续学习研究,包括非大语言模型单模态持续学习(Non-LLM Unimodal CL)、非大语言模型多模态持续学习(Non-LLM Multimodal CL)以及大语言模型中的持续学习(CL in LLM);第三部分深入分析当前MLLM持续学习研究状态,包括基准评估、架构创新及理论与实证研究总结;最后探讨该领域的挑战与未来方向,旨在推动相关技术发展。
原文摘要 · Abstract (English)
Recent advancements in Artificial Intelligence have led to the development of Multimodal Large Language Models (MLLMs). However, adapting these pre-trained models to dynamic data distributions and various tasks efficiently remains a challenge. Fine-tuning MLLMs for specific tasks often causes performance degradation in the model's prior knowledge domain, a problem known as 'Catastrophic Forgetting'. While this issue has been well-studied in the Continual Learning (CL) community, it presents new challenges for MLLMs. This review paper, the first of its kind in MLLM continual learning, presents an overview and analysis of 440 research papers in this area.The review is structured into four sections. First, it discusses the latest research on MLLMs, covering model innovations, benchmarks, and applications in various fields. Second, it categorizes and overviews the latest studies on continual learning, divided into three parts: non-large language models unimodal continual learning (Non-LLM Unimodal CL), non-large language models multimodal continual learning (Non-LLM Multimodal CL), and continual learning in large language models (CL in LLM). The third section provides a detailed analysis of the current state of MLLM continual learning research, including benchmark evaluations, architectural innovations, and a summary of theoretical and empirical studies.Finally, the paper discusses the challenges and future directions of continual learning in MLLMs, aiming to inspire future research and development in the field. This review connects the foundational concepts, theoretical insights, method innovations, and practical applications of continual learning for multimodal large models, providing a comprehensive understanding of the research progress and challenges in this field, aiming to inspire researchers in the field and promote the advancement of related technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。