解决视觉问答持续学习中的模态失衡问题
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
- 通过跨模态提示查询实现双模态平衡选择
- 迭代交互重建提示,提升知识保留率
- 适合长期多模态学习场景的模型优化
基于预训练模型的持续视觉问答(CVQA)通过提示调优实现了令人瞩目的进展。然而,现有方法多采用跨模态提示隔离策略,分别构建视觉与文本提示,加剧了模态不平衡,导致性能随时间下降。为此,我们提出MM-Prompt框架,包含跨模态提示查询与跨模态提示恢复机制。前者在提示查询阶段引入跨模态信号,实现更均衡的提示选择;后者通过迭代跨模态交互,结合对齐损失引导联合提示重构,防止表征漂移。大量实验表明,MM-Prompt在准确率与知识保留方面均优于现有方法,且在整个持续学习过程中保持双模态的均衡参与。
原文摘要 · Abstract (English)
Continual Visual Question Answering (CVQA) based on pre-trained models(PTMs) has achieved promising progress by leveraging prompt tuning to enable continual multi-modal learning. However, most existing methods adopt cross-modal prompt isolation, constructing visual and textual prompts separately, which exacerbates modality imbalance and leads to degraded performance over time. To tackle this issue, we propose MM-Prompt, a novel framework incorporating cross-modal prompt query and cross-modal prompt recovery. The former enables balanced prompt selection by incorporating cross-modal signals during query formation, while the latter promotes joint prompt reconstruction through iterative cross-modal interactions, guided by an alignment loss to prevent representational drift. Extensive experiments show that MM-Prompt surpasses prior approaches in accuracy and knowledge retention, while maintaining balanced modality engagement throughout continual learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。