arXiv:2607.02020cs.AI2026-07被引 1

模型答对题但依赖证据变了,新方法让答案和依据都稳定。

Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails

论文配图:Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails
图 1 · 摘自论文原文
  • 用反事实干预分析模型对视觉、文本等证据的依赖变化。
  • 在多个数据集上提升性能,减少证据依赖漂移和隐藏遗忘率。
  • 无需额外推理开销,适合持续学习的多模态应用。

多模态大语言模型需持续适应新任务与领域,但现有持续学习评估主要关注旧答案是否正确,忽视了多模态对齐的稳定性。本文研究这一被忽略的问题,探讨模型在持续学习中是否仍能保持正确的视觉、文本、OCR、图表及文档证据使用方式。我们发现存在‘隐性证据遗忘’现象:模型虽维持答案准确率,却悄然转向不同或更不相关的证据通道。为此提出无回放的依赖约束持续学习框架 RCL:冻结历史检查点作为行为参照,通过反事实通道干预估计教师与学生模型的证据依赖分布,并联合优化任务学习、预测保留与依赖保留,不增加推理成本。在 CoIN、COAST、MCITlib 及一个敏感证据的多模态流上,RCL 均显著优于无回放、PEFT、路由与记忆增强基线,大幅降低模态依赖漂移、主导证据翻转与隐藏遗忘率。结果表明,鲁棒的持续多模态学习需保留正确答案背后的证据路径,而不仅是答案本身。

原文摘要 · Abstract (English)

Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether old answers remain correct, leaving the stability of multimodal grounding largely unexamined. We study this overlooked failure mode and ask whether a continually adapted MLLM can preserve not only what it answers, but also how it uses visual, textual, OCR, chart, and document evidence. We identify \emph{hidden evidence-use forgetting}, where answer accuracy is retained while the model silently shifts toward different or less grounded evidence channels, and propose \textsc{RCL}, a replay-free reliance-constrained continual learning framework. \textsc{RCL} freezes the previous checkpoint as a behavioral reference, estimates teacher and student evidence-reliance profiles through counterfactual channel interventions, and jointly optimizes task learning, prediction preservation, and reliance preservation without adding inference-time cost. Across CoIN, COAST, MCITlib, and an evidence-sensitive multimodal stream, \textsc{RCL} consistently improves final performance and reduces forgetting over replay-free, PEFT, routing, and memory-assisted baselines, while substantially lowering modality reliance drift, dominant evidence flips, and hidden forgetting rates. These results suggest that robust continual multimodal learning requires preserving the evidence path behind correct answers, not merely the answers themselves.

多模态持续学习证据依赖模型稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。