arXiv:2603.10476cs.CLcs.AI2026-03被引 4

让大模型通过对话协商对齐集体价值,解决多利益相关者冲突。

Learning to Negotiate: Multi-Agent Deliberation for Collective Value Alignment in LLMs

  • 用对立角色自对弈生成协商对话,优化决策过程。
  • 在道德困境中达成共识能力提升,同时保持语言能力不下降。
  • 适合需要多方协调的AI系统,如政策制定、群体决策支持。

大模型对齐研究在单智能体场景中已取得进展,如基于人类反馈的强化学习(RLHF),近期也探索了基于AI反馈的强化学习(RLAIF)和动态对齐目标等可扩展方法。然而,这些方法在多利益相关者场景中仍受限,因存在价值观冲突,需开展协商性讨论。本文提出一种基于多智能体协商的对齐框架,使大模型对齐于集体能动性(Collective Agency, CA)这一已有对齐目标,同时增强冲突解决能力。为实现可扩展训练,两个自对弈的大模型被赋予对立人格,进行轮流对话以达成互利方案。我们生成合成的道德困境提示与冲突人格对,并使用外部大模型奖励模型,通过组相对策略优化(GRPO)进行策略优化。虽然奖励基于最终输出的CA得分,但梯度应用于对话令牌,以直接改进协商互动动态。实验表明,该模型在CA对齐效果上媲美单智能体基线,且显著提升冲突解决性能,同时未降低通用语言能力。结果表明,基于协商的推理训练为大模型在价值冲突场景下支持集体决策提供了可行路径。

原文摘要 · Abstract (English)

LLM alignment has progressed in single-agent settings through paradigms such as RL with human feedback (RLHF), while recent work explores scalable alternatives such as RL with AI feedback (RLAIF) and dynamic alignment objectives. However, these approaches remain limited in multi-stakeholder settings, where conflicting values arise and deliberative negotiation is required. This work proposes a multi-agent negotiation-based alignment framework that aligns LLMs to Collective Agency (CA)-an existing alignment objective introduced to promote the continual expansion of agency-while simultaneously improving conflict-resolution capability. To enable scalable training, two self-play LLM instances are assigned opposing personas and engage in turn-based dialogue to synthesize mutually beneficial solutions. We generate synthetic moral-dilemma prompts and conflicting persona pairs, and optimize the policy via RLAIF using Group Relative Policy Optimization (GRPO) with an external LLM reward model. While rewards are computed from CA scores assigned to the final completion, gradients are applied to dialogue tokens to directly improve deliberative interaction dynamics. Experiments show that the model achieves CA alignment comparable to a single-agent baseline while substantially improving conflict-resolution performance without degrading general language capabilities. These results suggest that negotiation-driven deliberation training provides a practical path toward LLMs that better support collective decision-making in value-conflict scenarios.

多智能体对齐协商大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。