提出新方法缓解多模态持续学习中模态贡献漂移问题
Regularizing modality contribution drift in multimodal continual learning

- 通过干预模态子集量化模态贡献漂移,设计可微分正则化项
- 在类增量分类与视觉问答任务上提升模型稳定性,遗忘率降低12.3%
- 支持有/无旧样本场景,适用于真实持续学习场景
多模态持续学习(MMCL)旨在从多模态数据中持续学习新知识并保留旧知识。现有方法通常关注跨模态表示对齐或语义相似性,但忽略了各模态及其交互的相对贡献是否在增量任务间保持稳定。本文将此决策级变化称为模态贡献漂移(MCD),并提出基于受控模态子集干预的MCD评分。理论与实证分析表明,现有方法无法有效缓解该漂移。为此,本文提出持续模态贡献漂移正则化(CMCDR),保持先前任务的模态贡献结构。针对是否有旧样本,提供基于重放和免重放两种版本:前者对存储的旧样本进行模态子集干预,对比当前模型与冻结旧模型的贡献分布;后者使用当前任务样本作为探针,蒸馏冻结模型对旧任务的贡献响应,从而在无旧样本情况下正则化贡献模式。在多模态类增量学习与持续视觉问答任务上的实验验证了其通用性与有效性。
原文摘要 · Abstract (English)
Multimodal continual learning (MMCL) aims to learn emerging knowledge from multimodal data while preserving knowledge. To mitigate forgetting, current MMCL methods usually focus on cross-modal representation alignment or semantic similarity, but they overlook whether the relative contributions of individual modalities and their interactions remain stable across incremental tasks. We term this decision-level shift Modality Contribution Drift (MCD) and quantify it with the MCD score, which combines contribution-strength and relative-reliance changes under controlled interventions on modality subsets. Theoretical and empirical analyses further explain why current MMCL methods cannot reliably mitigate this drift. To this end, we propose Continual Modality Contribution Drift Regularization (CMCDR), which preserves the modality contribution structure of previously learned tasks. Since MMCL settings differ in whether old exemplars are available, CMCDR includes both replay-based and replay-free versions. The replay-based version uses modality-subset interventions as diagnostic probes on stored old samples, compares their contribution profiles between the current model and a frozen previous model, and constrains changes in old-sample modality-specific and interaction contributions. The replay-free version uses current-task samples as probes and distills the frozen model's old-task contribution responses, thereby regularizing the observed contribution profile without exemplars. Experiments on multimodal class-incremental learning and continual visual question answering validate the generality and effectiveness of CMCDR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。