arXiv:2508.02076cs.AIcs.GT2025-08被引 4

用博弈论设计协作机制,让多个大模型主动贡献、避免搭便车。

Everyone Contributes! Incentivizing Strategic Cooperation in Multi-LLM Systems via Sequential Public Goods Games

  • 按顺序决策,前序输出影响后续贡献,形成激励相容的协作链。
  • 在推理、编程等任务上性能接近大模型,且显著优于单模型和传统协作方法。
  • 适合需要多模型协同、追求高效与公平合作的场景,如智能助手系统。

协调多个大语言模型(LLMs)协同完成复杂任务,在计算成本与集体性能之间存在根本权衡。我们提出一种基于博弈论的强化学习框架——多智能体协作顺序公共品博弈(MAC-SPGG),系统性激励多LLM集成中的合作行为。在MAC-SPGG中,LLM智能体按序行动,观察前序输出并更新信念以决定自身贡献。通过重构公共品奖励机制,努力付出成为唯一子博弈完美纳什均衡(SPNE),有效消除传统公共品博弈中的搭便车问题。其顺序协议替代高成本的轮次信息交换,实现通信开销降低的同时保持策略深度。我们在现实参数下证明了SPNE的存在性与唯一性,并实证表明,经MAC-SPGG训练的集成系统在推理、数学、代码生成及自然语言任务上超越单模型基线、思维链提示及其他协作方法,性能接近大规模模型。结果凸显了结构化、激励对齐的MAC-SPGG协作在可扩展、鲁棒的多智能体语言生成中的强大潜力。

原文摘要 · Abstract (English)

Coordinating multiple large language models (LLMs) to solve complex tasks collaboratively poses a fundamental trade-off between the computation costs and collective performance compared with individual model. We introduce a novel, game-theoretically grounded reinforcement learning (RL) framework, the Multi-Agent Cooperation Sequential Public Goods Game (MAC-SPGG), to systematically incentivize cooperation in multi-LLM ensembles. In MAC-SPGG, LLM agents move in sequence, observing predecessors' outputs and updating beliefs to condition their own contributions. By redesigning the public-goods reward, effortful contributions become the unique Subgame Perfect Nash Equilibrium (SPNE), which eliminates free-riding under traditional SPGG or PGG. Its sequential protocol replaces costly round-based information exchanges with a streamlined decision flow, cutting communication overhead while retaining strategic depth. We prove the existence and uniqueness of the SPNE under realistic parameters, and empirically show that MAC-SPGG-trained ensembles outperform single-agent baselines, chain-of-thought prompting, and other cooperative methods, even achieving comparable performance to large-scale models across reasoning, math, code generation, and NLP tasks. Our results highlight the power of structured, incentive-aligned MAC-SPGG cooperation for scalable and robust multi-agent language generation.

多模型协作博弈论激励机制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。