arXiv:2605.30698cs.CVcs.AI2026-05

让多个AI agents通过共享视觉证据达成可信共识

Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence

论文配图:Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence
图 1 · 摘自论文原文
  • 用视觉区域证据显式暴露各agent的推理依据
  • 通过证据一致性提升跨agent共识可靠性
  • 无需训练,适合实际部署且结果可解释

视觉语言模型(VLM)在视觉问答(VQA)任务中表现优异。为减少个体幻觉和盲区,多智能体协作通过汇聚多元视角成为有前景的范式。尽管文本问答中已取得成功,其在多模态领域的潜力仍待挖掘。现有方法多沿用文本中心协议,侧重文字讨论而忽略视觉信息对齐。本文揭示关键洞察:仅答案一致不足以保证可靠共识,必须依赖共享的视觉证据支持。为此,提出EAGLE(Evidence-Aligned Grounded Multi-agent Reasoning)框架,无需训练,通过显式暴露各agent的视觉定位区域作为证据,实现相互验证,并以证据一致性指导最终决策。在六个VQA基准上实验表明,EAGLE在跨领域平均性能最优,同时保持轻量、可解释且易于部署。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have achieved strong performance on visual question answering (VQA). To mitigate individual hallucinations and blind spots, aggregating diverse perspectives via multi-agent collaboration has emerged as a promising paradigm. While this approach has shown great success in textual QA, its potential in the multimodal domain remains under-explored. Existing multi-agent VQA methods predominantly adapt text-centric protocols, focusing on textual discussions while ignoring the alignment of visual information. In this work, we reveal a key insight: answer-level agreement is insufficient for reliable multi-agent VQA; \textit{aligned visual evidence} -- shared support from the image regions agents rely on -- is essential for trustworthy consensus. To leverage this insight, we propose EAGLE (\textbf{E}vidence-\textbf{A}ligned \textbf{G}rounded mu\textbf{L}ti-agent r\textbf{E}asoning), a training-free evidence-centered framework for coordinating multiple VLM agents. EAGLE explicitly exposes each agent's grounding regions as visual evidence, enables mutual verification over the evidence, and uses evidence consistency to guide final decision-making. Experiments on six VQA benchmarks show that EAGLE achieves best average performance across domains while remaining lightweight, interpretable, and practical for deployment.

多智能体视觉问答可解释性协同推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。