让大模型学会倾听伙伴意见,提升多人协作效率
Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- 设计新算法让模型主动吸收同伴干预信息
- 实验显示新模型促进共识达成,探索方案更多样
- 适合研究人机协作与安全对齐的学者参考
大型语言模型(LLMs)在代理场景中越来越多地作为人类合作者使用。因此,评估其在多轮、多主体任务中的协作能力变得愈发重要。本文基于人工智能对齐与安全中断文献,提出关于LLM驱动合作者与干预者之间协作行为的新理论见解。目标是学习一种理想的‘伙伴感知’合作者,通过智能收集同伴提供的干预信息,提升团队在任务相关命题上的共同认知(CG)对齐。我们发现,使用标准强化学习人类反馈(RLHF)等方法训练的模型往往忽视善意干预,导致群体共识难以建立。为此,我们采用双玩家改进动作马尔可夫决策过程(Modified-Action MDP)分析该次优行为,并提出中断式协作角色扮演者(ICR)——一种新型伙伴感知学习算法,以训练出最优共知对齐的合作者。在多个协作任务环境中测试表明,平均而言,ICR 更能促进成功的共识收敛,并探索更多样化的解决方案。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly being deployed in agentic settings where they act as collaborators with humans. Therefore, it is increasingly important to be able to evaluate their abilities to collaborate effectively in multi-turn, multi-party tasks. In this paper, we build on the AI alignment and safe interruptibility literature to offer novel theoretical insights on collaborative behavior between LLM-driven collaborator agents and an intervention agent. Our goal is to learn an ideal partner-aware collaborator that increases the group's common-ground (CG) alignment on task-relevant propositions-by intelligently collecting information provided in interventions by a partner agent. We show how LLM agents trained using standard RLHF and related approaches are naturally inclined to ignore possibly well-meaning interventions, which makes increasing group common ground non-trivial in this setting. We employ a two-player Modified-Action MDP to examine this suboptimal behavior of standard AI agents, and propose Interruptible Collaborative Roleplayer (ICR)-a novel partner-aware learning algorithm to train CG-optimal collaborators. Experiments on multiple collaborative task environments show that ICR, on average, is more capable of promoting successful CG convergence and exploring more diverse solutions in such tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。