arXiv:2501.09352cs.LGcs.MM2025-01被引 1

解决多模态增量学习中模态缺失问题,提升模型抗遗忘能力。

PAL: Prompting Analytic Learning with Missing Modality for Multi-Modal Class-Incremental Learning

  • 用特定模态提示补偿缺失信息,保持数据整体表征。
  • 将任务重构成递归最小二乘问题,获得解析解。
  • 无需样本存储,在模态缺失下仍保持高准确率。

多模态增量学习(MMCIL)旨在利用音视频、图像-文本等多模态数据,实现模型在连续任务中的持续学习并缓解遗忘。现有研究多关注多模态信息的融合与利用,但忽略了增量学习阶段模态缺失这一关键挑战,导致严重遗忘和性能下降。为此,我们提出PAL,一种面向模态缺失场景的新型无样本框架。具体而言,设计模态专属提示以补偿缺失信息,帮助模型维持数据的完整表征;在此基础上,将MMCIL重构为递归最小二乘问题,获得解析线性解。基于此,PAL不仅缓解了分析学习固有的欠拟合问题,还有效保留了缺失模态数据的整体表征,在多种多模态增量场景中显著减少遗忘,表现更优。大量实验表明,相比竞争方法,PAL在UPMC-Food101和N24News等多个数据集上均取得显著优势,展现出对模态缺失的鲁棒性及强抗遗忘能力。

原文摘要 · Abstract (English)

Multi-modal class-incremental learning (MMCIL) seeks to leverage multi-modal data, such as audio-visual and image-text pairs, thereby enabling models to learn continuously across a sequence of tasks while mitigating forgetting. While existing studies primarily focus on the integration and utilization of multi-modal information for MMCIL, a critical challenge remains: the issue of missing modalities during incremental learning phases. This oversight can exacerbate severe forgetting and significantly impair model performance. To bridge this gap, we propose PAL, a novel exemplar-free framework tailored to MMCIL under missing-modality scenarios. Concretely, we devise modality-specific prompts to compensate for missing information, facilitating the model to maintain a holistic representation of the data. On this foundation, we reformulate the MMCIL problem into a Recursive Least-Squares task, delivering an analytical linear solution. Building upon these, PAL not only alleviates the inherent under-fitting limitation in analytic learning but also preserves the holistic representation of missing-modality data, achieving superior performance with less forgetting across various multi-modal incremental scenarios. Extensive experiments demonstrate that PAL significantly outperforms competitive methods across various datasets, including UPMC-Food101 and N24News, showcasing its robustness towards modality absence and its anti-forgetting ability to maintain high incremental accuracy.

多模态增量学习模态缺失抗遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。