用多智能体分歧识别所需视觉工具,提升多模态推理准确率
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
- 通过智能体辩论分歧自动发现并调用图像检测、OCR等工具
- 在A-OKVQA和MMMU上分别比最强基线高3.4%和2.4%
- 可灵活适配新领域工具,医学数据集提升1.3%
专用视觉工具可为大语言模型或视觉语言模型引入专家知识(如定位、空间推理、医学知识等),但判断何时调用何种工具仍具挑战。我们提出DART,一种基于多个辩论视觉智能体分歧来识别有用工具(如目标检测、OCR、空间推理等)的多智能体框架。这些工具通过引入新信息促进智能体间有效讨论,并提供与工具对齐的共识得分,以突出与专家工具一致的智能体,从而推动讨论进程。我们采用聚合智能体结合各智能体输出与工具信息,选出最优答案。在四个多样化基准测试中验证,该方法优于多智能体辩论及单智能体工具调用框架,在A-OKVQA和MMMU上分别超越次强基线(带裁判模型的多智能体辩论)3.4%和2.4%。此外,DART在新应用领域中表现出良好适应性,在M3D医学数据集上相较其他强基线提升1.3%。我们还测量了多轮文本重叠度,凸显DART相较于现有方法更丰富的讨论内容;工具调用分布分析表明,多样工具被稳定用于解决分歧。
原文摘要 · Abstract (English)
Specialized visual tools can augment large language models or vision language models with expert knowledge (e.g., grounding, spatial reasoning, medical knowledge, etc.), but knowing which tools to call (and when to call them) can be challenging. We introduce DART, a multi-agent framework that uses disagreements between multiple debating visual agents to identify useful visual tools (e.g., object detection, OCR, spatial reasoning, etc.) that can resolve inter-agent disagreement. These tools allow for fruitful multi-agent discussion by introducing new information, and by providing tool-aligned agreement scores that highlight agents in agreement with expert tools, thereby facilitating discussion. We utilize an aggregator agent to select the best answer by providing the agent outputs and tool information. We test DART on four diverse benchmarks and show that our approach improves over multi-agent debate as well as over single agent tool-calling frameworks, beating the next-strongest baseline (multi-agent debate with a judge model) by 3.4% and 2.4% on A-OKVQA and MMMU respectively. We also find that DART adapts well to new tools in applied domains, with a 1.3% improvement on the M3D medical dataset over other strong tool-calling, single agent, and multi-agent baselines. Additionally, we measure text overlap across rounds to highlight the rich discussion in DART compared to existing multi-agent methods. Finally, we study the tool call distribution, finding that diverse tools are reliably used to help resolve disagreement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。