arXiv:2601.18735cs.AIcs.LG2026-01中稿 · ICLR被引 9

用市场机制交易视觉不确定性,让多智能体系统更省钱高效。

Why Keep Your Doubts to Yourself? Trading Visual Uncertainties in Multi-Agent Bandit Systems

  • 将认知不确定性拆解为可交易的资产,通过经济规则驱动协作。
  • 在5个基准上比最优基线高8.5%准确率,成本降低3倍以上。
  • 适合想构建低成本、可扩展视觉智能系统的研究者和开发者。

视觉语言模型(VLMs)推动了强大多智能体系统的发展,但其规模化面临经济不可持续问题:在信息不对称下协调异构智能体常导致成本飙升。现有范式如多智能体混合与基于知识的路由,依赖启发式代理,忽略成本且破坏不确定性结构,导致明显次优的协调结果。我们提出Agora框架,将协调重构为去中心化的不确定性市场。Agora将认知不确定性形式化为可结构化交易的资产(感知、语义、推理),并基于理性经济规则强制智能体间盈利驱动的交易。一个具备市场意识的代理,扩展了汤普森采样,发起协作并引导系统走向成本高效的均衡。在五个多模态基准(MMMU、MMBench、MathVision、InfoVQA、CC-OCR)上的实验表明,Agora优于强基线VLM和启发式多智能体策略,例如在MMMU上较最佳基线提升8.5%准确率,同时成本降低超3倍。这些结果确立了基于市场的协调是构建经济可行多智能体视觉智能系统的原理性与可扩展范式。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) enable powerful multi-agent systems, but scaling them is economically unsustainable: coordinating heterogeneous agents under information asymmetry often spirals costs. Existing paradigms, such as Mixture-of-Agents and knowledge-based routers, rely on heuristic proxies that ignore costs and collapse uncertainty structure, leading to provably suboptimal coordination. We introduce Agora, a framework that reframes coordination as a decentralized market for uncertainty. Agora formalizes epistemic uncertainty into a structured, tradable asset (perceptual, semantic, inferential), and enforces profitability-driven trading among agents based on rational economic rules. A market-aware broker, extending Thompson Sampling, initiates collaboration and guides the system toward cost-efficient equilibria. Experiments on five multimodal benchmarks (MMMU, MMBench, MathVision, InfoVQA, CC-OCR) show that Agora outperforms strong VLMs and heuristic multi-agent strategies, e.g., achieving +8.5% accuracy over the best baseline on MMMU while reducing cost by over 3x. These results establish market-based coordination as a principled and scalable paradigm for building economically viable multi-agent visual intelligence systems.

多智能体不确定性市场机制视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。