arXiv:2604.14779cs.CVcs.CL2026-04

针对视觉问答持续学习中的模型失衡问题,提出自适应掩码机制提升稳定性。

AIM: Asymmetric Information Masking for Visual Question Answering Continual Learning

论文配图:AIM: Asymmetric Information Masking for Visual Question Answering Continual Learning
图 1 · 摘自论文原文
  • 按模态敏感度对不同组件施加针对性掩码,平衡训练过程
  • 在VQA v2和GQA上实现最优平均性能与遗忘率
  • 特别适合需要保持组合推理能力的视觉语言模型

在持续视觉问答(Continual VQA)中,现有持续学习方法多基于对称、单模态架构设计,但现代视觉-语言模型(VLMs)的可训练组件天然具有非对称性。这种结构不匹配导致模型在持续学习时极易发生灾难性遗忘。具体而言,不对称性使标准全局正则化倾向于优化庞大的语言解码器,而较小但关键的视觉投影层则极易受到干扰。由此引发的局部退化严重削弱了模型的组合推理能力。为此,本文提出非对称信息掩码(AIM),通过基于模态特异性敏感度的靶向掩码,实现稳定与可塑性的平衡。在VQA v2和GQA上的持续学习实验表明,AIM在平均性能(AP)和平均遗忘率(AF)两项指标上均达到当前最优,同时更有效地保留了对新技能-概念组合的泛化能力。

原文摘要 · Abstract (English)

In continual visual question answering (VQA), existing Continual Learning (CL) methods are mostly built for symmetric, unimodal architectures. However, modern Vision-Language Models (VLMs) violate this assumption, as their trainable components are inherently asymmetric. This structural mismatch renders VLMs highly prone to catastrophic forgetting when learning from continuous data streams. Specifically, the asymmetry causes standard global regularization to favor the massive language decoder during optimization, leaving the smaller but critical visual projection layers highly vulnerable to interference. Consequently, this localized degradation leads to a severe loss of compositional reasoning capabilities. To address this, we propose Asymmetric Information Masking (AIM), which balances stability and plasticity by applying targeted masks based on modality-specific sensitivity. Experiments on VQA v2 and GQA under continual VQA settings show that AIM achieves state-of-the-art performance in both Average Performance (AP) and Average Forgetting (AF), while better preserving generalization to novel skill-concept compositions.

持续学习视觉问答多模态模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。