揭示视觉语言模型在持续学习中跨模态贡献的理论机制
Understanding Cross-Modal Contributions in Continual Vision-Language Models: A Theoretical Perspective

- 从理论角度分析视觉与语言模态在连续任务中的贡献变化
- 实验证明模型能有效捕捉不同任务间的跨模态依赖关系
- 适合研究持续学习与多模态模型的科研人员参考
持续视觉语言模型通常通过顺序微调来实现对新环境(任务)的适应,但这一范式在提升适应性的同时,会过度强调先前任务的贡献,损害已有知识的稳定性。尽管现有研究已关注视觉语言模型中的持续学习与灾难性遗忘问题,但对序列环境中模态特异性贡献的理论理解仍不充分。本文提出一种新的理论视角,用于分析连续任务中视觉与语言模态的跨模态贡献。我们在大型视觉语言模型上对理论发现进行实证评估,证明其能有效捕捉环境级别的跨模态贡献。分析揭示了模型在不同任务顺序和任务相似性下的贡献鲁棒性,并提升了泛化性能。
原文摘要 · Abstract (English)
Continual vision-language models are commonly addressed through sequential fine-tuning; however, although this paradigm enables adaptation to new environments (tasks), it inherently emphasizes the contribution of previously learned environments (tasks) at the expense of the stability required to preserve previously acquired knowledge. While existing approaches have adequately studied continual learning and catastrophic forgetting in vision-language models (VLMs), the theoretical understanding of modality-specific contributions across a sequence of environments remains largely unexplored. In this paper, we present a new theoretical perspective to understand the cross-modal (vision-language) contributions to consecutive environments. We empirically evaluate our theoretical findings on large VLMs and demonstrate their effectiveness in capturing environment-level cross-modal contributions. Our analysis provides deeper insights into continual VLMs, highlighting their contribution robustness to varying task orders and inter-task similarities, and their improved generalization performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。