arXiv:2512.03318cs.AI2025-12NeurIPS被引 17

测试大模型代理在复杂社交情境下的通用合作能力

Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia

  • 用Concordia环境评估代理在零样本混合动机场景中的协作能力
  • 实测显示当前代理在说服与规范执行任务中表现不足
  • 适合关注AI社交智能与通用协作能力的研究者参考

大型语言模型(LLM)代理在社交互动中展现出令人印象深刻的能力,正越来越多地部署于需与人类及人工代理交互的场景。这类交互是LLM代理的关键前沿领域,但现有评估方法无法衡量其在新社交情境中的泛化能力。本文提出一种基于Concordia——一个自然语言多智能体模拟环境——的评估方法,用于检验LLM代理在零样本、混合动机环境中实现合作的能力。该方法通过测试代理在多样化伙伴与情境中识别并利用互利机会的能力,衡量其通用合作智能。我们在NeurIPS 2024 Concordia竞赛中呈现了实证结果,评估了代理在从谈判到集体行动问题等多样场景中达成互利成果的表现。研究发现,当前代理能力与实现可靠协作所需的稳健泛化之间存在显著差距,尤其在需要说服力与规范执行的情境中表现更弱。

原文摘要 · Abstract (English)

Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing evaluation methods fail to measure how well these capabilities generalize to novel social situations. In this paper, we introduce a method for evaluating the ability of LLM-based agents to cooperate in zero-shot, mixed-motive environments using Concordia, a natural language multi-agent simulation environment. Our method measures general cooperative intelligence by testing an agent's ability to identify and exploit opportunities for mutual gain across diverse partners and contexts. We present empirical results from the NeurIPS 2024 Concordia Contest, where agents were evaluated on their ability to achieve mutual gains across a suite of diverse scenarios ranging from negotiation to collective action problems. Our findings reveal significant gaps between current agent capabilities and the robust generalization required for reliable cooperation, particularly in scenarios demanding persuasion and norm enforcement.

大模型代理社交智能协作评估Concordia

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。