受大脑海马体启发,提升视频语言模型持续学习能力
Bisecle: Binding and Separation in Continual Learning for Video Language Understanding
- 借鉴海马体的绑定与分离机制,设计跨模态监督与对比提示学习
- 在多个VideoQA基准上显著减少遗忘,提升跨任务泛化能力
- 适合需要长期适应新视频场景的智能系统开发者
前沿视觉语言模型在视频理解任务中表现优异。然而,现实世界视频通常是连续演化的数据流(如可穿戴眼镜捕捉的动态场景),要求模型持续适应变化的数据分布和新场景。由于微调计算成本过高,通常仅更新少量参数而冻结大部分模型,这对大型多模态基础模型的持续学习框架带来新挑战:灾难性遗忘与更新冲突。现有方法难以应对参数高效持续学习问题。人类海马体则演化出高效的记忆形成与巩固机制。受此启发,本文提出Bisecle,用于视频语言持续学习:通过多方向监督模块捕捉更多跨模态关系,并设计对比提示学习方案隔离任务特定知识,以实现高效记忆存储。绑定与分离过程增强模型对复杂经验的保留能力,支持在视频理解任务中稳健高效的持续学习。我们在多个VideoQA基准上进行充分评估,验证了Bisecle在缓解遗忘、提升跨任务泛化方面的有效性。
原文摘要 · Abstract (English)
Frontier vision-language models (VLMs) have made remarkable improvements in video understanding tasks. However, real-world videos typically exist as continuously evolving data streams (e.g., dynamic scenes captured by wearable glasses), necessitating models to continually adapt to shifting data distributions and novel scenarios. Considering the prohibitive computational costs of fine-tuning models on new tasks, usually, a small subset of parameters is updated while the bulk of the model remains frozen. This poses new challenges to existing continual learning frameworks in the context of large multimodal foundation models, i.e., catastrophic forgetting and update conflict. While the foundation models struggle with parameter-efficient continual learning, the hippocampus in the human brain has evolved highly efficient mechanisms for memory formation and consolidation. Inspired by the rapid Binding and pattern separation mechanisms in the hippocampus, in this work, we propose Bisecle for video-language continual learning, where a multi-directional supervision module is used to capture more cross-modal relationships and a contrastive prompt learning scheme is designed to isolate task-specific knowledge to facilitate efficient memory storage. Binding and separation processes further strengthen the ability of VLMs to retain complex experiences, enabling robust and efficient continual learning in video understanding tasks. We perform a thorough evaluation of the proposed Bisecle, demonstrating its ability to mitigate forgetting and enhance cross-task generalization on several VideoQA benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。