arXiv:2602.05302cs.AI2026-02被引 1

用真实商业谈判场景评估大模型协商能力,发现顶尖模型表现已超人类。

PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios

  • 构建多智能体真实谈判基准PieArena,覆盖三种交互模式。
  • 前沿模型(如GPT-5)在谈判结果上达到甚至超过人类专家水平。
  • 不仅看成交结果,还分析模型行为偏差与诚信度,适合评估智能体可靠性。

我们深入评估大语言模型在协商这一核心商业任务中的表现,该任务需要策略推理、心理理论和价值创造能力。为此,我们提出PieArena——一个基于精英商学院MBA课程真实谈判场景的大型多智能体交互基准。评估涵盖镜像对战、交叉对战和人-模型对战三种配对方式。我们开发了一种连续谈判收益的排名模型,实现无序不变、不确定性量化,并校正实验系统性偏差。进一步研究联合意图型智能体支架的影响,发现中低层级模型获益显著,而前沿模型收益递减。作为校准基准,我们采集了受训商学院学生的真人对战及人-模型谈判数据,结果显示代表性前沿语言模型(GPT-5)在评估设置下表现匹配或超越人类基准。除交易结果外,PieArena提供多维行为画像,揭示模型在指令遵循、计算准确性、裁判评估的欺骗行为与声誉方面的异质性,凸显超越单一结果评价的价值。

原文摘要 · Abstract (English)

We present an in-depth evaluation of LLMs' ability to negotiate, a central business task requiring strategic reasoning, theory of mind, and economic value creation. To do so, we introduce PieArena, a large-scale negotiation benchmark grounded in multi-agent interactions over realistic scenarios adapted from MBA negotiation courses at an elite business school. We evaluate language agents across three pairing regimes: mirror-play, cross-play, and human-LM play. We develop a ranking model for continuous negotiation payoffs that yields order-invariant, uncertainty-quantified leaderboards while correcting for systematic experimental asymmetries. We further study the effects of joint-intentionality agentic scaffolding and find asymmetric gains, with large improvements for mid- and lower-tier LMs and diminishing returns for frontier LMs. As calibration anchors, we collect human-human and human-LM negotiation data from trained business school students, finding that a representative frontier language agent (GPT-5) matches or exceeds this human baseline in our evaluation settings. Beyond deal outcomes, PieArena provides a multi-dimensional behavioral profile that reveals cross-model heterogeneity in instruction compliance, computation accuracy, as well as judge-assessed deception and reputation, illustrating the value of evaluation beyond outcome-only leaderboards.

大模型评估谈判模拟多智能体行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。