提出UNIFIER框架,让多模态大模型在不同视觉场景中持续学习不遗忘。
Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives
- 通过视觉表征扩展和一致性约束,实现跨场景知识共享
- 在20步连续学习中,问答准确率提升2.7%~10.6%,F1提升3.4%~7.7%
- 适用于户外、水下、低空、室内等多场景设备端视觉任务
部署在设备上的多模态大语言模型(MLLMs)需适应背景与视角不断变化的视觉场景,以完成复杂视觉任务。为研究真实场景切换下的灾难性遗忘问题,我们构建了覆盖高空、水下、低空和室内四种环境的多模态视觉理解数据集MSVQA。同时提出UNIFIER(mUltimodal coNtInual learning with MLLMs From multi-scenarIo pERspectives),一种针对视觉差异的持续学习框架。相比现有方法,UNIFIER通过视觉表征扩展(VRE)和视觉一致性约束(VCC),在同场景中积累知识,并实现跨场景相互增强。实验表明,在20步跨场景持续学习任务中,其最终阶段的VQA得分提升2.70%~10.62%,F1得分提升3.40%~7.69%,优于当前最优方法QUAD。MSVQA数据集已公开于https://huggingface.co/datasets/Kaij00/MSVQA。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) deployed on devices must adapt to continuously changing visual scenarios such as variations in background and perspective, to effectively perform complex visual tasks. To investigate catastrophic forgetting under real-world scenario shifts, we construct a multimodal visual understanding dataset (MSVQA), covering four distinct scenarios and perspectives: high-altitude, underwater, low-altitude, and indoor environments. Furthermore, we propose UNIFIER (mUltimodal coNtInual learning with MLLMs From multi-scenarIo pERspectives), a continual learning (CL) framework designed to address visual discrepancies while learning different scenarios. Compared to existing CL methods, UNIFIER enables knowledge accumulation within the same scenario and mutual enhancement across different scenarios via Vision Representation Expansion (VRE) and Vision Consistency Constraint (VCC). Experimental results show that UNIFIER improves the last-step VQA scores by 2.70%~10.62% and the last-step F1 scores by 3.40%~7.69% compared to the state-of-the-art method, QUAD, in 20-step cross-scenario continual learning tasks. MSVQA dataset is available at https://huggingface.co/datasets/Kaij00/MSVQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。