测试大模型团队中角色替换的代价,发现任务表现影响小但协作效率下降显著。
Testing Interchangeability in LLM Agent Teams
- 通过角色互换实验,检验多智能体系统中代理的可替代性。
- 互换后任务得分损失小,但单位进展通信量增加16%至63%。
- 长期协作团队互换代价更大,适合关注协作稳定性的研究者。
生产环境中的多智能体系统常频繁更换角色代理,基于代理可互换的假设。本文对此进行验证:在相同任务下,从同一基础模型生成八支独立团队,每名代理在十轮组建中保持私有笔记;随后交换角色匹配的代理,并评估其在保留任务上的表现。与仅模拟阵容变动但不更换代理的对照组相比,代理互换导致任务得分损失较小,但单位进展所需通信量上升16%至63%;在汉牌(Hanabi)中,被替换的代理表现甚至不如新手,暗示与前搭档形成的协作惯例产生干扰。在协作烹饪(Collab-Overcooked)中,若主导议程的代理被替换,大部分新增通信来自未更换的代理。三种消融实验(基础模型、解码温度、组建长度)显示,互换惩罚与独立组建团队间的偏差程度一致:贪婪解码降低惩罚,团队历史翻倍则提升惩罚。在这些设置中,代理在任务结果上更具可替代性,但在协调效率上差异明显,且组建历史越长,互换代价越大。
原文摘要 · Abstract (English)
Production multi-agent systems replace agents constantly, on the assumption that an agent filling a role is interchangeable with any other agent that can do the job. We test that assumption. Eight teams per setting are formed independently from one base model on the same tasks, each agent keeping a private notebook across ten formation episodes; we then trade role-matched agents between teams and measure what changes on held-out tasks. Against a placebo that reproduces the disruption of a roster change without changing who occupies the seat, a swap costs little in task score but raises the communication a team spends per unit of progress by 16 to 63 percent, and in Hanabi a swapped agent is more expensive than an inexperienced one, consistent with interference from conventions learned with its former partner. In Collab-Overcooked, when the agent that sets the agenda is replaced, most of the extra communication comes from the agent that stayed. Three ablations, over base models, decoding temperature and formation length, move the swap penalty alongside one other quantity: how far independently formed teams drift apart. Greedy decoding lowers both; doubling a team's history raises both. In these settings, agents are more fungible in task outcome than in coordination efficiency, with larger swap effects after longer formation histories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。