用预训练模型实现多模态增量学习,缓解遗忘并提升融合效果
Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion
- 基于专家混合结构的增量特征提取器,支持多模态微调
- 自适应融合模块提升特征判别力,文本多样性增强策略有效
- 新设计对比损失与评估指标,适用于多模态增量场景
传统多模态增量学习(MCIL)仅关注视觉与文本,本文拓展至视觉、音频与文本三模态,解决信息互补整合难与灾难性遗忘问题。提出基于多模态预训练模型的MCIL方法:首先引入基于混合专家(MoE)结构的多模态增量特征提取器(MIFE),实现AudioCLIP的有效增量微调;其次设计自适应音视频融合模块(AAVFM),包含掩码阈值机制与动态特征融合机制,并提出文本多样性增强策略;第三,提出新型多模态增量对比训练损失,优化跨模态对齐;最后引入两个专用于MCIL的评估指标。在三个多模态数据集上进行大量实验,验证了方法的有效性。
原文摘要 · Abstract (English)
Unlike traditional Multimodal Class-Incremental Learning (MCIL) methods that focus only on vision and text, this paper explores MCIL across vision, audio and text modalities, addressing challenges in integrating complementary information and mitigating catastrophic forgetting. To tackle these issues, we propose an MCIL method based on multimodal pre-trained models. Firstly, a Multimodal Incremental Feature Extractor (MIFE) based on Mixture-of-Experts (MoE) structure is introduced to achieve effective incremental fine-tuning for AudioCLIP. Secondly, to enhance feature discriminability and generalization, we propose an Adaptive Audio-Visual Fusion Module (AAVFM) that includes a masking threshold mechanism and a dynamic feature fusion mechanism, along with a strategy to enhance text diversity. Thirdly, a novel multimodal class-incremental contrastive training loss is proposed to optimize cross-modal alignment in MCIL. Finally, two MCIL-specific evaluation metrics are introduced for comprehensive assessment. Extensive experiments on three multimodal datasets validate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。