提出MacVQA框架,提升持续视觉问答的记牢与适应能力
MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering
- 自适应记忆分配+全局噪声过滤,优化特征表示
- 平均准确率43.38%,遗忘率仅2.32%(标准任务)
- 适合需要长期学习的视觉问答系统
视觉问答(VQA)需融合多模态信息进行推理。随着持续学习的发展,知识保留与新信息适应已取得进展,但现有方法仍难平衡知识保持、适应性与鲁棒特征表示。为此,本文提出新型框架MacVQA,结合自适应记忆分配与全局噪声过滤,融合视觉与问题信息并抑制噪声,采用基于原型的记忆分配机制优化特征质量与内存使用。该设计使模型在持续VQA学习中实现知识获取、保留与组合泛化间的良好平衡。在10个持续VQA任务上的实验表明,MacVQA在标准任务上达到43.38%平均准确率和2.32%平均遗忘率,在新组合任务上达42.53%平均准确率和3.60%平均遗忘率,优于现有基线。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) requires models to reason over multimodal information, combining visual and textual data. With the development of continual learning, significant progress has been made in retaining knowledge and adapting to new information in the VQA domain. However, current methods often struggle with balancing knowledge retention, adaptation, and robust feature representation. To address these challenges, we propose a novel framework with adaptive memory allocation and global noise filtering called MacVQA for visual question answering. MacVQA fuses visual and question information while filtering noise to ensure robust representations, and employs prototype-based memory allocation to optimize feature quality and memory usage. These designs enable MacVQA to balance knowledge acquisition, retention, and compositional generalization in continual VQA learning. Experiments on ten continual VQA tasks show that MacVQA outperforms existing baselines, achieving 43.38% average accuracy and 2.32% average forgetting on standard tasks, and 42.53% average accuracy and 3.60% average forgetting on novel composition tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。