arXiv:2605.27586cs.MAcs.CL2026-05

一个对齐代理可让其他未训练代理自发合作,效果显著提升。

You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents

论文配图:You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
图 1 · 摘自论文原文
  • 用教师模型的对话能力训练种子代理,通过语言影响他人。
  • 合作率从24.8%升至62.2%,在新环境中零样本达成91.5%交易成功率。
  • 适合研究多智能体协同与高效对齐策略的学者参考。

在分布式开放多智能体系统中,确保智能体行为对齐极具挑战,尤其当群体规模扩大且存在非对齐代理时。本文发现,仅需一个对齐代理,即可通过自然语言交互将合作行为传播给未训练的代理,这一现象称为对齐传播。我们在红黑博弈(Red-Black Game)中验证:该任务是基于团队的迭代囚徒困境,队友需讨论并投票决定集体行动。通过将教师模型的协作推理与说服性对话提炼至Qwen-3-14B,生成种子代理;将其置于四个未训练队友中,合作率从24.8%提升至62.2%,优于教师模型及原始Gemini-3.1-Pro。更惊人的是,该种子代理仅在红黑博弈中训练,却在空间化生存模拟Sugarscape中实现零样本迁移,达到91.5%的交易成功率,远超21.6%的基线。结果表明,多智能体对齐可从逐个训练转向可通过策略性放置种子代理实现的可扩展社交能力。

原文摘要 · Abstract (English)

Ensuring agent behaviors in distributed open multi-agent systems remains challenging, especially as populations grow and unaligned agents may exist. We show that a single aligned agent can propagate cooperative behaviors to untrained agents purely through natural language interaction, a phenomenon we term Alignment Propagation. We study this in the Red-Black Game, a team-based iterated Prisoner's Dilemma in which teammates deliberate and vote to determine their team's collective action. By distilling the cooperative reasoning and persuasive dialogues of a teacher model into a Qwen-3-14B, we obtain a seed agent that, when placed among four untrained teammates, doubles the cooperation rate from 24.8% to 62.2%, outperforming the teacher model and a vanilla Gemini-3.1-Pro. Remarkably, a seed trained exclusively on the RedBlack Game transfers zero-shot to Sugarscape, a spatially grounded survival simulation with pairwise trading, achieving a 91.5% trade success rate versus a 21.6% baseline. Our results reframe multi-agent alignment from an exhaustive per-agent training problem to a scalable social capability that can be engineered through strategic seed placement.

多智能体对齐传播零样本迁移语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。