用博弈论框架让大模型谈判失败原因可诊断,超越单纯看成交率。
TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate

- 构建贝叶斯博弈环境,显式建模对手隐藏状态与策略
- 13个前沿模型在成交率饱和下仍暴露策略短板
- 可定位具体失败环节,适合优化谈判类AI系统
谈判是经济交换的核心机制,也是评估智能体语言模型的经典场景,需在隐藏偏好、策略沟通和约束条件下进行多轮交互。由于缺乏内在验证器,现有评估依赖大模型间对战或成交率等整体指标,难以揭示失败原因。本文提出Terms-Bench(多轮策略中的经济推理测试平台),采用贝叶斯博弈框架,将对手的隐含类型、策略与收益结构显式设定,使环境本身成为验证器。在双边价格谈判场景中,对手的私有状态和模拟策略对智能体隐藏,但对评估者可见。这使对手从黑箱对手转变为诊断工具,实现失败归因和最优性差距分析。对13个来自主流厂商的前沿大模型评估显示:尽管成交率已饱和,模型在剩余价值获取、线索利用、信念校准和合规性上差异显著,暴露出此前基准无法捕捉的个体化谈判瓶颈。
原文摘要 · Abstract (English)
Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation. It is also a canonical testbed for agentic language models, requiring multi-turn interaction under hidden preferences, strategic communication, and binding constraints. These properties make negotiation hard to evaluate: unlike math or code, it has no intrinsic verifier. Existing LLM negotiation evaluations rely on LLM-vs.-LLM interaction or aggregate outcomes such as deal rate, leaving failures opaque. We introduce Terms-Bench, short for Testbed for Economic Reasoning in Multi-turn Strategy, a Bayesian-game framework that makes the environment itself the verifier by specifying the counterpart's latent type, policy, and payoff structure. We instantiate it in bilateral price negotiation, where the counterpart's private state and simulator policy are hidden from the agent but observable to the evaluator. This turns the counterpart from a black-box opponent into a diagnostic instrument, enabling agent-attributable failure analysis and oracle-reference optimality gaps. Evaluating 13 LLM agents spanning frontier systems from major providers, Terms-Bench turns negotiation evaluation from aggregate ranking into actionable diagnosis: where agents fail, why they fail, and what to strengthen. Empirically, frontier models saturate deal rate yet diverge in surplus extraction, cue use, belief calibration, and compliance, revealing agent-specific bargaining bottlenecks masked by prior benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。