arXiv:2509.00798cs.CVcs.AI2025-09ACL被引 2

通过分步检索与推理,提升视觉问答中知识获取与融合的准确性。

Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering

  • 分阶段检索多源知识,动态调整查询策略以增强信息覆盖。
  • 在六个基准上平均提升1.8%的问答准确率,召回率显著改善。
  • 适合需要精准知识融合的复杂视觉问答任务,如科研或医疗场景。

知识密集型视觉问答(VQA)需要超出图像内容的外部知识,要求精确的视觉定位和视觉与文本信息的连贯融合。尽管多模态检索增强生成已取得显著进展,但现有方法多采用单次流程,常无法获取充分知识,且缺乏修正错误推理的机制。我们提出PMSR(渐进式多模态搜索与推理),通过构建结构化推理轨迹,增强知识获取与整合能力。PMSR使用双范围查询,基于最新记录和推理路径从异构知识库中检索多样化知识,并通过组合推理将证据合成紧凑记录。该设计支持可控的迭代优化,减少错误传播,实现更稳定的推理过程。在六个不同基准(Encyclopedic-VQA、InfoSeek、MMSearch、LiveVQA、FVQA 和 OK-VQA)上的广泛实验表明,PMSR持续提升检索召回率和端到端答案准确率。

原文摘要 · Abstract (English)

Knowledge-intensive visual question answering (VQA) requires external knowledge beyond image content, demanding precise visual grounding and coherent integration of visual and textual information. Although multimodal retrieval-augmented generation has achieved notable advances by incorporating external knowledge bases, existing approaches largely adopt single-pass frameworks that often fail to acquire sufficient knowledge and lack mechanisms to revise misdirected reasoning. We propose PMSR (Progressive Multimodal Search and Reasoning), a framework that progressively constructs a structured reasoning trajectory to enhance both knowledge acquisition and synthesis. PMSR uses dual-scope queries conditioned on both the latest record and the trajectory to retrieve diverse knowledge from heterogeneous knowledge bases. The retrieved evidence is then synthesized into compact records via compositional reasoning. This design facilitates controlled iterative refinement, which supports more stable reasoning trajectories with reduced error propagation. Extensive experiments across six diverse benchmarks (Encyclopedic-VQA, InfoSeek, MMSearch, LiveVQA, FVQA, and OK-VQA) demonstrate that PMSR consistently improves both retrieval recall and end-to-end answer accuracy.

视觉问答多模态知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。