测试大模型在供应链谈判中如何分配利益,发现模型能力、身份和提示设计决定谈判成败。
When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains

- 用大模型模拟买卖双方谈判,对比其效率与收益分配。
- 模型平均耗时2.98轮,损失21%-34%潜在收益,部分模型会接受亏损合同。
- 模型所属厂商影响分钱比例,提示词可操控谈判结果,适合研究AI决策者的人看。
随着大语言模型从辅助决策转向自主采购,企业需了解委托谈判代理是否创造价值、能否稳定分配利益,以及避免亏损合约。本文研究经典供应链谈判问题:买家拥有私有需求信息,与不了解情况的卖家协商数量-价格合同。我们对OpenAI、Google和阿里巴巴的九个大模型进行了基准测试,对比经验证的完美贝叶斯均衡,在9,840次大模型间谈判中评估表现。首先,能力决定价值创造:模型在98.9%的谈判中达成协议,捕获未折现的首优盈余95.4%,但平均需2.98轮,而基准仅1.25轮,此延迟导致21%-34%盈余损失;能力也影响可靠性:基础模型在19.2%情况下接受个体非理性合约,而中高端模型仅0.0%-0.6%,故自动化利润验证是低阶模型的必要防线。其次,盈余分配具有关系性:提供方身份比能力排名更能预测谁获得更多利益——自对战买家中,OpenAI均分40%,Google为50%,阿里Qwen达70%,该排序在受限沟通与无折现条件下仍成立。调换供应商角色可使分配变动7-18个百分点,且能力强的Qwen旗舰反而是跨家族中最弱的卖方,说明选择供应商是关键分配决策。第三,提示词是战略杠杆:委托机制将委托人经济耐心与代理人提示后的策略耐心分离,这一免费部署选择成为盈余分配的最强驱动因素(解释了90%方差)。综合三方面,本文建立了一个基于均衡的审计框架,评估大模型在折扣效率、分配特征与操作可靠性三个维度的表现。
原文摘要 · Abstract (English)
As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed seller. We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations. First, capability governs value creation. Agents agree in 98.9% of negotiations and capture 95.4% of first-best surplus undiscounted, but average 2.98 rounds against the benchmark's 1.25, and this delay erodes 21-34% of surplus. Capability also governs reliability: baseline models accept individually irrational contracts in 19.2% of cases, versus 0.0-0.6% at mid-tier and flagship, making automated profit verification the binding guardrail below that threshold. Second, surplus capture is relational. Provider identity predicts who captures surplus better than capability rank: self-play buyer shares average 40% for OpenAI, 50% for Google, and 70% for Alibaba's Qwen, an ordering that survives restricted communication and no discounting. Reversing which provider sells moves the division by 7-18 percentage points, and the capable Qwen flagship is the weakest cross-family seller: vendor choice is a first-order distributional decision. Third, the prompt is a strategic lever. Delegation separates the principal's economic patience from the agent's prompted strategic patience, a free deployment choice that is the single strongest driver of surplus division (90% of explained variance). Together these establish an equilibrium-referenced audit of AI agents along three dimensions: discounted efficiency, distributional profile, and operational reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。