让多模态智能体辩论更可靠,提升视觉语言推理准确率。
WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
- 分角色设计求解者与验证者,动态加权反馈结果。
- 在多个数据集上比现有方法提升2-7%准确率。
- 适合需要多智能体协作的复杂视觉语言任务研究者。
当前大语言模型因训练数据多样而具备互补优势,多智能体辩论(MAD)被用于强化推理能力,但主要局限于纯语言任务,对多模态问题的适用性仍不明确。本文研究将MAD应用于视觉-语言推理问题,提出一种通用且模块化的框架WISE,将智能体分为求解者(Solvers)和验证者(Reflectors),前者生成解答,后者验证正确性、分配权重并提供自然语言反馈。为融合多轮辩论结果,我们改进了Dawid-Skene算法,实现对响应差异与权重的联合建模。在SMART-840、VisualPuzzles、EvoChart-QA及新构建的SMART-840++数据集(程序生成、难度可控)上评估,结果显示,WISE在多种多模态任务和LLM配置下,相较最先进的MAD设置与聚合方法,准确率持续提升2-7%。
原文摘要 · Abstract (English)
Recent large language models (LLMs) are trained on diverse corpora and tasks, leading them to develop complementary strengths. Multi-agent debate (MAD) has emerged as a popular way to leverage these strengths for robust reasoning, though it has mostly been applied to language-only tasks, leaving its efficacy on multimodal problems underexplored. In this paper, we study MAD for solving vision-and-language reasoning problems. Our setup enables generalizing the debate protocol with heterogeneous experts that possess single- and multi-modal capabilities. To this end, we present Weighted Iterative Society-of-Experts (WISE), a generalized and modular MAD framework that partitions the agents into Solvers, that generate solutions, and Reflectors, that verify correctness, assign weights, and provide natural language feedback. To aggregate the agents' solutions across debate rounds, while accounting for variance in their responses and the feedback weights, we present a modified Dawid-Skene algorithm for post-processing that integrates our two-stage debate model. We evaluate WISE on SMART-840, VisualPuzzles, EvoChart-QA, and a new SMART-840++ dataset with programmatically generated problem instances of controlled difficulty. Our results show that WISE consistently improves accuracy by 2-7% over the state-of-the-art MAD setups and aggregation methods across diverse multimodal tasks and LLM configurations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。