用强化学习让大模型学会多买家谈判中的策略平衡。
Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations

- 通过可验证奖励训练,让模型学会在探索与谈判间权衡。
- 相比前沿模型,谈判盈余提升显著,且能稳定签约高价值买家。
- 策略对未知买家风格和预算分布具有强泛化能力。
谈判是管理科学中的基础战略互动,涉及各方在保护私有信息(如保留成本、隐藏估值)的同时达成协议。典型复杂场景是单一卖家同时与多个拥有异质私有预算的买家谈判,受限于有限的沟通回合,卖家需在广泛市场探索以发现最高估值与集中精力争取单个买家之间取得平衡。我们分析发现,主流大语言模型虽语言流畅,但作为经济决策者表现不佳:缺乏市场探索意识,常固守当前最高报价,忽视潜在高估值买家。本文提出基于可验证奖励的强化学习(RLVR)训练方法,将奖励函数锚定于客观经济结果,使市场发现与利润提取之间的战略平衡自然涌现。实验表明,训练后的卖家经历多阶段策略演化,掌握价格锚定与战略性探询技巧,显著提升谈判盈余,不仅增强说服力,还能持续锁定高价值对手。最终,该策略在未见过的买家谈判风格与预算分布下仍具鲁棒泛化能力。
原文摘要 · Abstract (English)
Negotiation is a fundamental strategic interaction in management science, characterized by agents attempting to reach agreements while protecting private information, such as reservation costs and hidden valuations. A prevalent yet complex scenario involves a single seller negotiating concurrently with multiple buyers, each possessing heterogeneous, private budgets. In such settings, constrained by a limited number of communication turns, the seller must balance exploring the broader market to discover the highest valuation with concentrating sufficient turns on a single target buyer to secure the best possible outcome. Our analysis reveals a significant gap in standard Large Language Models (LLMs): while these models are linguistically proficient, they fail to act as effective economic decision-makers. Specifically, they exhibit a failure to explore the buyer pool, often fixating on the current highest bid rather than strategically investigating the market to discover latent high valuations. In this paper, we propose a specialized training recipe using Reinforcement Learning from Verifiable Rewards (RLVR). By anchoring the reward function to objective economic outcomes, the strategic balance between market discovery and surplus extraction emerges natively through the learning process. Our results demonstrate that the trained seller undergoes a multi-stage strategic evolution, learning to leverage price anchoring and strategic probing to identify more profitable counterparties. The agent extracts a substantially higher surplus than frontier models by both improving its persuasive bargaining skills and consistently closing deals with high-value buyers. Finally, we show that our seller strategies generalize robustly to unseen buyer negotiation styles and budget distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。