无需额外提示,仅凭答案就能学会组合推理的视觉问答新方法
Disentanglement-Based Equivariant Learning for Compositional VQA

- 用因果干预解耦视觉语言概念,重构表示
- 通过等变约束增强模型对新组合的推理能力
- 适合需要强泛化能力的视觉问答场景
组合式视觉问答(Compositional VQA)要求模型理解先前未见的概念组合,是极具挑战性的任务。现有方法常忽略概念解耦,且依赖额外训练信号,难以在真实场景应用。本文提出基于解耦的等变学习框架(DEAL),仅使用真实答案进行训练。DEAL通过因果启发的干预,在重编码框架中解耦视觉与文本输入中的概念;基于等变性原则,对推理输入进行组合变换,并对输出施加等变约束,以增强模型的组合推理能力。在基准数据集CLEVR-CoGenT和GQA-SGL上的大量实验表明,DEAL在视觉与语言泛化设置下均优于现有最先进方法。
原文摘要 · Abstract (English)
Compositional visual question answering (VQA) represents a challenging yet fundamental task that requires models to comprehend novel combinations of previously learned concepts. The current methods often overlook the disentanglement of underlying concepts and are restricted in terms of their ability to effectively capture the compositional variation mechanism. Moreover, the state-of-the-art techniques depend on additional clues for training, which is not feasible in real-world VQA scenarios. To address these issues, in this paper, we introduce a novel Disentanglement-based EquivAriant Learning (DEAL) framework for compositional VQA, which is guided exclusively by ground-truth answers. In DEAL, we employ causality-inspired interventions to disentangle concepts derived from visual and textual inputs within a re-encoding framework. Based on the principle of equivariance, we subsequently perform a compositional transformation on the inference input and impose the equivariant constraint on the output to augment the compositional reasoning capacity of the model. Comprehensive experiments conducted on the benchmark CLEVR-CoGenT and GQA-SGL datasets validate the superiority of our proposed DEAL approach over the existing state-of-the-art methods for compositional VQA tasks in both visual and linguistic generalization settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。