多智能体协作解谜,用博弈论提升视觉语言推理准确率
GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning
- 将推理拆成感知与验证两类智能体,通过博弈机制协同工作
- 小模型提升5-6%,大模型如GPT-4o也增2-3%准确率
- 支持可解释推理,适合需要可靠决策的复杂视觉任务
我们提出GAM-Agent,一种基于博弈论的多智能体框架,用于增强视觉语言推理。不同于以往单智能体或整体模型,GAM-Agent将推理过程建模为基础智能体(负责视觉子任务)与关键验证智能体之间的非零和博弈,后者确保逻辑一致性和事实正确性。智能体间通过结构化陈述、证据及不确定性估计进行通信。框架引入不确定性感知控制器,可在意见分歧或模糊时触发多轮辩论,从而生成更鲁棒且可解释的预测。在四个挑战性基准(MMMU、MMBench、MVBench、V*Bench)上的实验表明,GAM-Agent显著提升了各类视觉语言模型的表现。值得注意的是,其对中小型模型(如Qwen2.5-VL-7B、InternVL3-14B)的准确率提升达5–6%,对强模型如GPT-4o也有最高2–3%的增益。该方法模块化、可扩展且通用,为实现可靠、可解释的多智能体多模态推理提供了新路径。
原文摘要 · Abstract (English)
We propose GAM-Agent, a game-theoretic multi-agent framework for enhancing vision-language reasoning. Unlike prior single-agent or monolithic models, GAM-Agent formulates the reasoning process as a non-zero-sum game between base agents--each specializing in visual perception subtasks--and a critical agent that verifies logic consistency and factual correctness. Agents communicate via structured claims, evidence, and uncertainty estimates. The framework introduces an uncertainty-aware controller to dynamically adjust agent collaboration, triggering multi-round debates when disagreement or ambiguity is detected. This process yields more robust and interpretable predictions. Experiments on four challenging benchmarks--MMMU, MMBench, MVBench, and V*Bench--demonstrate that GAM-Agent significantly improves performance across various VLM backbones. Notably, GAM-Agent boosts the accuracy of small-to-mid scale models (e.g., Qwen2.5-VL-7B, InternVL3-14B) by 5--6\%, and still enhances strong models like GPT-4o by up to 2--3\%. Our approach is modular, scalable, and generalizable, offering a path toward reliable and explainable multi-agent multimodal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。