arXiv:2608.26095cs.CVcs.AI2026-08

提出视觉依赖感知框架,让多模态模型持续学习新数据时既不遗忘旧知识又高效适应新任务。

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

论文配图:A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training
图 1 · 摘自论文原文
  • 基于视觉依赖结构失真设计对抗性迁移机制,防止跨模态遗忘
  • 利用视觉依赖异质性增强新任务学习,提升模型适应能力
  • 适合需要长期演进的多模态大模型部署场景

本文提出多模态无监督持续后训练(MU-CPT)新任务,使已部署的多模态大模型能从流式无标签数据中持续演化。现有方法对目标词元统一优化,忽视其词元级视觉依赖(VD)的异质性。我们揭示:视觉依赖结构失真可指示跨模态灾难性遗忘,其内在异质性则可作为新任务学习的导航。为此,提出视觉依赖感知(VDA)框架:一是视觉约束最优传输(VC-OT),将旧任务视觉依赖失真建模为最优传输问题,通过区域感知基代价与依赖分层传输惩罚,防止视觉焦点全局偏移并抑制视觉依赖退化为语言偏差;二是视觉调制适应(VMA),利用视觉依赖异质性强化视觉锚定的新任务学习,促进新任务可塑性。实验在自建MU-CPT设置下验证了VDA在保持旧任务稳定性的同时实现新任务高效学习的有效性。

原文摘要 · Abstract (English)

In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Existing unsupervised post-training methods for MLLMs typically optimize target tokens uniformly, overlooking their heterogeneous visual dependence (VD). However, we reveal that token-level VD is crucial for MU-CPT. Specifically, its structural distortion serves as an indicator of cross-modal catastrophic forgetting, and its inherent heterogeneity acts as a compass to guide new-task learning. Leveraging this property, we propose a Visual Dependence-Aware (VDA) framework with two main components. First, Visually Constrained Optimal Transport (VC-OT) formulates the VD structural distortion of old-task VD during new-task learning as an optimal transport problem to mitigate cross-modal forgetting. By designing a region-aware ground cost and a dependence-stratified transport penalty, it prevents global shifts in visual focus while strictly prohibiting visual reliance from degenerating into language bias. Second, Visually Modulated Adaptation (VMA) exploits VD heterogeneity to emphasize visually grounded new-task learning, promoting new-task plasticity. Together, our method simultaneously maintains old-task stability and new-task plasticity during challenging MU-CPT. Extensive experiments under our MU-CPT setting validate the effectiveness of VDA.

多模态持续学习视觉依赖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。