提出多视角推理框架,让模型既能看全局又能抓关键对象。
See the Forest and the Trees: A Synergistic Reasoning Framework for Knowledge-Based Visual Question Answering
- 并行生成全局、结构和因果三类证据,实现多角度推理
- 在OK-VQA等三个数据集上刷新最佳性能,提升显著
- 可无缝接入现有模型,方法设计比模型规模更重要
多模态大语言模型(MLLM)推动了基于知识的视觉问答(KBVQA)的发展,但其推理仍受限于单一维度的证据。这种‘只见树木不见森林’的方式难以实现全面理解。受‘既见森林也见树木’理念启发,我们提出Synergos-VQA——一种协同推理框架。该框架在推理时同步生成并融合三类互补证据:(1)整体证据,用于感知整个场景(森林);(2)结构证据,通过原型驱动模块识别关键物体(树木);(3)因果证据,借助反事实探针确保推理稳健性。通过协同融合这些多维证据,模型实现更全面可靠的推理。大量实验表明,Synergos-VQA在三个挑战性基准(包括OK-VQA和A-OKVQA)上确立新基准。此外,该方法具备强即插即用能力,显著提升多种开源MLLM表现,证明优秀的方法设计可超越单纯模型规模。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have pushed the frontiers of Knowledge-Based Visual Question Answering (KBVQA), yet their reasoning is fundamentally bottlenecked by a reliance on uni-dimensional evidence. This "seeing only the trees, but not the forest" approach prevents robust, multi-faceted understanding. Inspired by the principle of seeing both the forest and trees, we propose Synergos-VQA, a novel synergistic reasoning framework. At its core, Synergos-VQA concurrently generates and fuses three complementary evidence streams at inference time: (1) Holistic Evidence to perceive the entire scene (the "forest"), (2) Structural Evidence from a prototype-driven module to identify key objects (the "trees"), and (3) Causal Evidence from a counterfactual probe to ensure the reasoning is robustly grounded. By synergistically fusing this multi-faceted evidence, our framework achieves a more comprehensive and reliable reasoning process. Extensive experiments show that Synergos-VQA decisively establishes a new state-of-the-art on three challenging benchmarks, including OK-VQA and A-OKVQA. Furthermore, our approach demonstrates strong plug-and-play capabilities, significantly boosting various open-source MLLMs and proving that superior methodological design can outperform sheer model scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。