用博弈论设计多智能体优先级策略,让小模型也能媲美大模型。
GARL: Game-Theoretic Reinforcement Learning for Multi-Agent Strategic Prioritisation

- 将优先级分配建模为两阶段博弈,通过角色化奖励信号优化策略
- 在法律争议点排序任务中,小模型性能超越闭源大模型
- 适合需要战略决策与多智能体协作的场景,如法律、政策分析
基于大语言模型的多智能体系统在战略决策任务中日益重要。其表现不仅取决于单个模型能力,更依赖于智能体间的交互策略。现有多智能体强化学习的奖励设计常任务特定且缺乏对交互结构的强约束。为此,我们提出GARL——一种用于多智能体战略优先级的博弈论强化学习框架。GARL将战略优先级建模为两阶段博弈:竞争性智能体首先在共享候选集上分配战略资源,随后高层仲裁者生成最终排名。由此产生的博弈论效用被转化为角色特异性强化信号,指导策略优化。我们在争议点排序任务中验证了GARL,结果表明其提升了排序性能,使小型开源LLM在相同候选集设定下可与强闭源模型竞争,并在法律领域能力和广义战略决策上取得提升。整体而言,GARL展示了如何将博弈论交互结构转化为强化学习目标,为多智能体战略优先级提供了一种原则性优化方法。
原文摘要 · Abstract (English)
LLM-based multi-agent systems are increasingly used for strategic decision-making tasks. In such settings, performance depends not only on individual model capabilities, but also on the policies by which agents interact and adapt. Multi-agent reinforcement learning can optimise these interaction policies, but its reward design often remains task-specific and weakly grounded in interaction structure. To address this gap, we propose GARL, a GAme-theoretic Reinforcement Learning framework for multi-agent strategic prioritisation. GARL formalises strategic prioritisation as a two-stage game: competing agents first allocate strategic resources over a shared candidate set, and a higher-level arbiter then produces the final ranking. The resulting game-theoretic utilities are converted into role-specific reinforcement signals, allowing policy optimisation to be guided by structured interaction. We instantiate GARL on issues-in-dispute ranking, where the goal is to prioritise core issues in legal proceedings. Experiments show that GARL improves ranking performance, enables small open-source LLMs to become competitive with a strong closed-source LLM under the same candidate-ranking setting, and yields gains in legal-domain competence and broader strategic decision-making. Overall, GARL demonstrates how game-theoretic interaction structure can be turned into reinforcement-learning objectives, providing a principled approach to policy optimisation in multi-agent strategic prioritisation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。