arXiv:2604.09855cs.AIcs.CL2026-04被引 4

用可验证奖励教大模型谈判,小模型胜过十倍大的对手。

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards

论文配图:Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards
图 1 · 摘自论文原文
  • 用经济盈余和预算约束做奖励信号,训练模型谈判。
  • 30B模型在抽取盈余上超越十倍大的前沿模型。
  • 能应对未知强对手和敌对卖家,泛化能力强。

大型语言模型(LLMs)虽具备自主交互潜力,但在不完全信息博弈如双边价格谈判中表现不佳。本文研究基于可验证奖励的强化学习(RLVR)能否有效教导LLM谈判。我们构建框架,让一个中等规模买家代理与受控的LLM卖家在真实商品分布上进行对抗训练。通过直接以最大化经济盈余和严格遵守私有预算为奖励信号,揭示出一种全新的四阶段战略演化:从盲目讨价还价,到使用激进起价,经历僵局期,最终发展出高阶说服能力。结果表明,该可验证训练使30B模型在盈余提取上显著优于十倍大小的前沿模型。此外,训练后的代理对训练中未见的更强对手具有鲁棒泛化能力,并在面对敌意卖家角色时仍保持有效性。

原文摘要 · Abstract (English)

The recent advancement of Large Language Models (LLMs) has established their potential as autonomous interactive agents. However, they often struggle in strategic games of incomplete information, such as bilateral price negotiation. In this paper, we investigate if Reinforcement Learning from Verifiable Rewards (RLVR) can effectively teach LLMs to negotiate. Specifically, we explore the strategic behaviors that emerge during the learning process. We introduce a framework that trains a mid-sized buyer agent against a regulated LLM seller across a wide distribution of real-world products. By grounding reward signals directly in the maximization of economic surplus and strict adherence to private budget constraints, we reveal a novel four-phase strategic evolution. The agent progresses from naive bargaining to using aggressive starting prices, moves through a phase of deadlock, and ultimately develops sophisticated persuasive skills. Our results demonstrate that this verifiable training allows a 30B agent to significantly outperform frontier models over ten times its size in extracting surplus. Furthermore, the trained agent generalizes robustly to stronger counterparties unseen during training and remains effective even when facing hostile, adversarial seller personas.

大模型强化学习谈判可验证奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。