arXiv:2503.14481cs.LGcs.CL2025-03被引 15

让AI通过协作自玩学会自我认知,判断何时该用工具、何时该沉默。

Don't lie to your friends: Learning what you know from collaborative self-play

  • 多智能体协作中通过集体奖励机制,自发形成对自身能力的认知。
  • 在异构工具环境下,协作群体策略可迁移到单个智能体并提升工具使用效率。
  • 适合研究自主智能体决策与可信推理的开发者或研究人员。

为成为有用的助手,人工智能代理必须了解自身的能与不能,包括何时从参数知识回答、何时使用工具、何时信任工具输出、何时回避或保留。这些能力难以通过监督微调教授,因为需要反映代理具体能力的示例。为此,我们提出一种全新方法:协同自玩(collaborative self-play)。构建多智能体协作系统,群体因共同得出正确答案而获得奖励。所需元知识从互动结构内置的激励中自然涌现。我们聚焦由拥有异构工具(如特定语料库检索)的小型智能体组成的社会,必须协作以最大化成功并最小化努力。实验表明,群体级奖励可诱导出可迁移的策略,在个体部署时显著改善工具使用与选择性预测性能。

原文摘要 · Abstract (English)

To be helpful assistants, AI agents must be aware of their own capabilities and limitations. This includes knowing when to answer from parametric knowledge versus using tools, when to trust tool outputs, and when to abstain or hedge. Such capabilities are hard to teach through supervised fine-tuning because they require constructing examples that reflect the agent's specific capabilities. We therefore propose a radically new approach to teaching agents what they know: \emph{collaborative self-play}. We construct multi-agent collaborations in which the group is rewarded for collectively arriving at correct answers. The desired meta-knowledge emerges from the incentives built into the structure of the interaction. We focus on small societies of agents that have access to heterogeneous tools (corpus-specific retrieval), and therefore must collaborate to maximize their success while minimizing their effort. Experiments show that group-level rewards for multi-agent communities can induce policies that \emph{transfer} to improve tool use and selective prediction in settings where individual agents are deployed in isolation.

多智能体自玩元认知工具使用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。