arXiv:2509.03817cs.AIcs.MA2025-09AAAI被引 14

让大模型团队学会自我反思,动态调整合作策略。

Learning to Deliberate: Meta-policy Collaboration for Agentic LLMs with Multi-agent Reinforcement Learning

  • 设计可自主决策的元策略,支持持续、优化与让步动作
  • 在5个推理任务上准确率提升4-5%,超越6种主流方法
  • 适合需要自适应协作的复杂推理场景

大型语言模型的多智能体系统在复杂推理中展现潜力,但其效果常受限于固定的协作协议。现有框架多关注宏观调度,忽视智能体内部的元认知能力。本文提出元策略反思框架(MPDF),让智能体学习一组高层元认知动作:持续、优化与让步。为解决传统策略梯度在此类设置中的不稳定性,提出SoftRankPO强化学习算法,通过平滑正态分位数映射奖励排名来塑造优势,增强训练对奖励波动的鲁棒性。实验表明,采用SoftRankPO的MPDF在五个数学与通用推理基准上平均准确率相较六种先进启发式与学习型多智能体推理方法提升4-5个百分点。本工作提出了学习自适应元认知策略的新范式,将重点从设计固定协议转向学习动态反思策略。

原文摘要 · Abstract (English)

Multi-agent systems of large language models (LLMs) show promise for complex reasoning, but their effectiveness is often limited by fixed collaboration protocols. These frameworks typically focus on macro-level orchestration while overlooking agents' internal deliberative capabilities. This critical meta-cognitive blindspot treats agents as passive executors unable to adapt their strategy based on internal cognitive states like uncertainty or confidence. We introduce the Meta-Policy Deliberation Framework (MPDF), where agents learn a decentralized policy over a set of high-level meta-cognitive actions: Persist, Refine, and Concede. To overcome the instability of traditional policy gradients in this setting, we develop SoftRankPO, a novel reinforcement learning algorithm. SoftRankPO stabilizes training by shaping advantages based on the rank of rewards mapped through smooth normal quantiles, making the learning process robust to reward variance. Experiments show that MPDF with SoftRankPO achieves a a 4-5% absolute gain in average accuracy across five mathematical and general reasoning benchmarks compared to six state-of-the-art heuristic and learning-based multi-agent reasoning algorithms. Our work presents a paradigm for learning adaptive, meta-cognitive policies for multi-agent LLM systems, shifting the focus from designing fixed protocols to learning dynamic, deliberative strategies.

多智能体元认知强化学习推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。