首次系统梳理视觉语言模型持续学习的挑战与解决方案。
Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- 提出四类应对遗忘的范式:多模态重放、跨模态正则化、参数高效适配与模型融合。
- 发现生成型模型存在'对齐税',导致深度思维链推理系统性崩溃。
- 适合研究持续学习、多模态大模型及智能体系统的学者参考。
视觉语言模型(VLMs)从预测型到生成式多模态大语言模型(MLLMs),通过强大的跨模态对齐与零样本泛化能力革新了人工智能。然而,在非平稳数据下实现持续学习仍面临重大挑战,其跨模态对齐与泛化能力极易受灾难性遗忘影响。与传统单模态持续学习不同,VLMs 面临跨模态特征漂移、共享架构引起的参数干扰以及零样本能力退化等独特问题。此外,生成式 MLLMs 出现独特的“对齐税”,表现为不仅记忆丢失,更导致深层思维链(CoT)推理系统性崩溃。本综述首次系统诊断预测型 VLM 与生成式 MLLMs 的持续学习问题,提出以挑战驱动的四类范式:(1) 多模态重放策略应对显式与隐式记忆漂移;(2) 跨模态正则化强化拓扑与几何对齐;(3) 参数高效适配利用动态路由与子空间投影;(4) 新兴的模型融合与解耦范式。我们批判性分析评估协议演进,强调双轨基准(领域与能力持续学习)的重要性。最后,勾勒未来研究路线图,聚焦组合式零样本学习、具身智能与传感器融合、自主代理生态系统。所有资源见:https://github.com/YuyangSunshine/Awesome-Continual-learning-of-Vision-Language-Models
原文摘要 · Abstract (English)
Vision-language models (VLMs), spanning predictive architectures to generative Multimodal Large Language Models (MLLMs), have revolutionized artificial intelligence through powerful cross-modal alignment and zero-shot generalization. However, enabling them to learn continually from non-stationary data remains a major challenge, as their cross-modal alignment and generalization capabilities are particularly vulnerable to catastrophic forgetting. Unlike traditional unimodal continual learning (CL), VLMs face unique challenges such as cross-modal feature drift, parameter interference due to shared architectures, and zero-shot capability erosion. Furthermore, generative MLLMs exhibit a unique "alignment tax," where catastrophic forgetting manifests not merely as factual amnesia, but as a systemic collapse of deep Chain-of-Thought (CoT) reasoning. This survey presents the first comprehensive diagnostic review bridging continual learning across predictive VLMs and generative MLLMs. We systematically deconstruct the aforementioned failure modes and propose a challenge-driven taxonomy comprising four core paradigms: (1) Multi-Modal Replay Strategies addressing explicit and implicit memory drift; (2) Cross-Modal Regularization enforcing topological and geometric alignment; (3) Parameter-Efficient Adaptation utilizing dynamic routing and subspace projections; and the emerging (4) Model Fusion and Decoupling paradigms. We critically analyze the evolution of evaluation protocols, highlighting the essential shift toward dual-track benchmarks (Domain vs. Ability CL). Finally, we chart a roadmap for future research, emphasizing compositional zero-shot learning, embodied AI with sensor fusion, and autonomous agentic ecosystems. All resources are available at: https://github.com/YuyangSunshine/Awesome-Continual-learning-of-Vision-Language-Models
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。