arXiv:2604.16472cs.GTcs.AI2026-04被引 3

用私有信息下的双边谈判测试大模型博弈能力,发现策略需精准定价与耐心等待。

Training Language Models for Bilateral Trade with Private Information

  • 设计事件驱动模拟器,用工具调用分离报价与语言交流,支持自动化评估。
  • 强模型按物品价值比例调整策略,弱模型仅在宽议价区间表现良好。
  • 训练中监督微调提升收益但降低成交率,强化学习可恢复成交率但损失收益。

不完全信息下的双边谈判为评估大语言模型智能体能力提供了受控实验场景。双边贸易要求个体理性、战略盈余最大化及合作以实现交易收益。我们构建了一个结构化谈判环境,允许大模型通过工具调用在事件驱动的模拟器中协商,将约束性报价与自然语言消息分离,实现自动化评估。该环境兼具双重用途:作为前沿模型的基准测试平台,也作为开放权重模型通过强化学习进行训练的环境。基准实验中,五种前沿模型(共15,000次谈判)的循环赛显示,有效策略通过连续报价实施价格歧视;激进锚定、校准让步与时间耐心与最高盈余份额和成交率正相关。被动妥协型策略在买方角色中削弱价格歧视,导致最低盈余捕获与成交完成率。更强模型能按物品价值比例扩展行为,在不同价格层级保持性能;较弱模型仅在广泛可能议价区间内表现良好。训练实验中,我们对Qwen3(8B、14B)采用监督微调(SFT)后接组相对策略优化(GRPO),对抗固定前沿对手。两阶段分别优化竞争目标:SFT使盈余份额约翻倍但降低成交率,强化学习恢复成交率却侵蚀盈余优势,反映奖励结构影响。SFT还压缩了各价格层级间的盈余差异,该特性泛化至未见对手,表明行为克隆培养的是比例策略而非记忆特定价格点。

原文摘要 · Abstract (English)

Bilateral bargaining under incomplete information provides a controlled testbed for evaluating large language model (LLM) agent capabilities. Bilateral trade demands individual rationality, strategic surplus maximization, and cooperation to realize gains from trade. We develop a structured bargaining environment where LLMs negotiate via tool calls within an event-driven simulator, separating binding offers from natural-language messages to enable automated evaluation. The environment serves two purposes: as a benchmark for frontier models and as a training environment for open-weight models via reinforcement learning. In benchmark experiments, a round-robin tournament among five frontier models (15,000 negotiations) reveals that effective strategies implement price discrimination through sequential offers. Aggressive anchoring, calibrated concession, and temporal patience correlate with the highest surplus share and deal rate. Accommodating strategies that concede quickly disable price discrimination in the buyer role, yielding the lowest surplus capture and deal completion. Stronger models scale their behavior proportionally to item value, maintaining performance across price tiers; weaker models perform well only when wide zones of possible agreement offset suboptimal strategies. In training experiments, we fine-tune Qwen3 (8B, 14B) via supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) against a fixed frontier opponent. These stages optimize competing objectives: SFT approximately doubles surplus share but reduces deal rates, while RL recovers deal rates but erodes surplus gains, reflecting the reward structure. SFT also compresses surplus variation across price tiers, which generalizes to unseen opponents, suggesting that behavioral cloning instills proportional strategies rather than memorized price points.

博弈论语言模型强化学习谈判模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。