arXiv:2509.05882cs.CLcs.AI2025-09中稿 · as an extended abs…被引 6

不同对齐方法影响AI协作效果,强化推理可提升多智能体合作成功率。

Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes

  • 设计角色扮演模拟,测试多种对齐方式在多轮协作中的表现。
  • 能抗动作修改的干预代理使任务正确率显著提升。
  • 适合研究AI协作、对齐机制或多智能体系统的设计者。

随着大语言模型(LLMs)被广泛集成到各类工作流中,它们正被视为与人类及其他AI系统协同工作的伙伴。若这些AI协作者需可靠地与人类或其他人工智能系统协调行为,其在多轮交互中的属性和表现必须可预测。本文研究不同对齐方法如何影响LLM智能体在多轮、多方协作中的表现。通过引入干预代理,该代理不提供答案,而是促使群体放慢节奏并反思推理过程以促进审慎决策。传统对齐技术多基于简化单用户场景,假设底层令牌马尔可夫决策过程(token MDP)最优。借助改进动作马尔可夫决策过程(modified-action MDP)的理论视角,我们发现这些方法忽视了长周期多主体交互动态。本文提出一种新型角色扮演模拟方法:按不同对齐方式训练LLM,并部署于协作任务对话中,量化干预对群体协作轨迹、信念一致性及协调性的影响。结果表明,在支持正确任务结果方面,对动作修改具有鲁棒性的干预代理显著优于常见对齐基线。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) get integrated into diverse workflows, they are increasingly being regarded as "collaborators" with humans, and required to work in coordination with other AI systems. If such AI collaborators are to reliably coordinate their actions and behaviors with humans or other AIs, their properties and behaviors over multi-turn interactions must be known and predictable. This paper examines how different alignment methods affect LLM agents' effectiveness as partners in multi-turn, multi-party collaborations. We study this question through the lens of intervention agents that insert themselves into group dialogues not to provide answers, but to encourage the collaborative group to slow down and reflect upon their reasoning for deliberative decision-making. Common alignment techniques are typically developed under simplified single-user settings and assume the optimality of the underlying token MDP. Using the theoretical lens of the modified-action MDP, we show how they do not account for the dynamics of long-horizon multi-party interactions. We present a novel roleplay simulation methodology, where we align LLMs according to different methods and then deploy them in collaborative task dialogues to quantify how interventions affect the trajectory of group collaboration, belief alignment, and coordination. Our results show that an intervention agent that is robust to action modification significantly outperforms common alignment baselines in supporting correct task outcomes.

多智能体对齐方法协作推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。