研究大模型推理对谈判表现的影响及成本代价
The Price of Thought: A Multilingual Analysis of Reasoning, Performance, and Cost of Negotiation in Large Language Models
- 通过自对弈实验评估推理能力对谈判效果的提升
- 推理使GPT-5性能提升31.4%,但计算成本增加近400%
- 开源模型常用英语内省推理,影响多语言可解释性
谈判是智能体面临的核心挑战,需具备战略推理、对手建模以及合作与竞争平衡能力。本文首次系统评估显式推理训练对商业与开源大模型谈判能力的影响,对比其在三种语言下的表现。基于三种对话游戏的自对弈设置,分析了性能与成本权衡、跨语言推理一致性及策略适应性。结果表明,启用推理(即扩展推理期计算)显著改善谈判结果,增强协作并克服任务复杂性,但带来巨大计算开销:推理使GPT-5性能提升31.4%,成本上升近400%。最关键发现是存在显著多语言推理差异:开源模型在德语或意大利语谈判中,内部推理步骤始终切换至英语,可能削弱推理轨迹披露带来的可解释性优势;而领先商业模型则保持推理与输出的语言一致。
原文摘要 · Abstract (English)
Negotiation is a fundamental challenge for AI agents, as it requires an ability to reason strategically, model opponents, and balance cooperation with competition. We present the first comprehensive study that systematically evaluates how explicit reasoning training affects the negotiation abilities of both commercial and open-weight large language models, comparing these models to their vanilla counterparts across three languages. Using a self-play setup across three diverse dialogue games, we analyse trade-offs between performance and cost, the language consistency of reasoning processes, and the nature of strategic adaptation exhibited by models. Our findings show that enabling reasoning -- that is, scaling test time compute -- significantly improves negotiation outcomes by enhancing collaboration and helping models overcome task complexities, but comes at a substantial computational cost: reasoning improves GPT-5's performance by 31.4 % while increasing its cost by nearly 400 %. Most critically, we uncover a significant multilingual reasoning distinction: open-weight models consistently switch to English for their internal reasoning steps, even when negotiating in German or Italian (and thus possibly impacting potential explainability gains through the disclosure of reasoning traces), while a leading commercial model maintains language consistency between reasoning and final output.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。