新数据集+去偏框架,让视听问答更鲁棒。
FortisAVQA and MAVEN: a Benchmark Dataset and Debiasing Framework for Robust Multimodal Reasoning
- 构建新数据集FortisAVQA,增强问题多样性并引入分布变化
- 提出MAVEN框架,通过多路径协同去偏使准确率提升7.81%
- 可插拔集成到多种模型,适合做多模态鲁棒性研究的团队
视听问答(AVQA)是一项需基于音视频配对输入回答自然语言问题的挑战性任务。现有方法常因过度适应数据集偏差而表现欠佳。为此,我们提出新数据集FortisAVQA,分两阶段构建:(1) 重写MUSIC-AVQA测试集的问题以扩大测试空间;(2) 引入问题分布偏移,实现对稀有、常见及总体分布下模型鲁棒性的精细评估。同时,提出鲁棒的多模态视听认知网络MAVEN,采用多面协同去偏策略减少偏差学习。实验表明,该架构在FortisAVQA上达到当前最优性能,准确率提升7.81%。消融实验证明各去偏组件有效。评估还揭示现有方法鲁棒性有限。此外,我们的策略可跨模型、跨数据集无缝集成,具有良好的可扩展性。代码与数据集已开源。
原文摘要 · Abstract (English)
Audio-Visual Question Answering (AVQA) is a challenging multimodal reasoning task requiring intelligent systems to answer natural language queries based on paired audio-video inputs accurately. However, existing AVQA approaches often suffer from overfitting to dataset biases, leading to poor robustness. Moreover, current datasets may not effectively diagnose these methods. To address these challenges, we first introduce a novel dataset, FortisAVQA, constructed in two stages: (1) rephrasing questions in the test split of the public MUSIC-AVQA dataset and (2) introducing distribution shifts across questions. The first stage expands the test space with greater diversity, while the second enables a refined robustness evaluation across rare, frequent, and overall question distributions. Second, we introduce a robust Multimodal Audio-Visual Epistemic Network (MAVEN) that leverages a multifaceted cycle collaborative debiasing strategy to mitigate bias learning. Experimental results demonstrate that our architecture achieves state-of-the-art performance on FortisAVQA, with a notable improvement of 7.81\%. Extensive ablation studies on both datasets validate the effectiveness of our debiasing components. Additionally, our evaluation reveals the limited robustness of existing multimodal QA methods. We also verify the plug-and-play capability of our strategy by integrating it with various baseline models across both datasets. Our dataset and code are available at https://github.com/reml-group/fortisavqa.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。