arXiv:2605.12789cs.RO2026-05

提升视觉语言模型持续学习能力,减少遗忘并保持跨模态对齐。

Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention

论文配图:Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
图 1 · 摘自论文原文
  • 用改进的EWC结合多模态信息矩阵和自适应正则化。
  • 遗忘率降低78%,跨模态对齐基本不变,仅增15%计算开销。
  • 适合需动态学习的机器人、自动驾驶等实时应用系统。

视觉语言大模型(如CLIP、Flamingo、BLIP)在图像描述、视觉问答和跨模态检索等任务中表现优异。但在顺序学习新任务时面临灾难性遗忘,尤其在多模态场景中,保持跨模态对齐进一步增加了学习难度。本文提出一种针对视觉语言模型的持续学习框架,融合增强版弹性权重固化(EWC)与参数高效微调技术,引入多模态费舍尔信息矩阵计算、模态间一致性保持机制及考虑视觉与文本编码器依赖关系的自适应正则化。实验表明,该框架相较朴素顺序训练,遗忘率降低78%,在序列学习过程中保持了跨模态对齐,额外计算成本仅为15%。本工作推动了多模态持续学习的前沿进展,可直接应用于自动驾驶、智能机器人助手及需持续学习的自适应机器人系统。

原文摘要 · Abstract (English)

Large language-vision models (LVLMs) such as CLIP, Flamingo, and BLIP have revolutionized AI by enabling understanding across textual and visual modalities. These models excel at tasks like image captioning, visual question answering, and cross-modal retrieval. However, they face catastrophic forgetting when learning new tasks sequentially, particularly challenging in multi-modal settings where preserving cross-modal alignments adds complexity to the learning process. This paper presents a comprehensive continual learning framework for LVLMs that combines enhanced Elastic Weight Consolidation (EWC) with parameter-efficient fine-tuning techniques. We integrate multi-modal Fisher Information Matrix calculation, consistency preservation across modalities, and adaptive regularization that considers dependencies across visual and textual encoders. The framework achieves a 78% reduction in forgetting rates relative to naive sequential training approaches through extensive evaluation testing. The framework also preserves alignment between modalities during sequential learning with only 15% additional computational cost. This work advances the state of the art in lifelong learning for multi-modal AI systems, with direct applications to autonomous driving, intelligent robotic assistants, and adaptive robotic systems that must continuously learn in dynamic real-world environments.

持续学习视觉语言模型跨模态对齐灾难性遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。