arXiv:2512.13063cs.CLcs.AI2025-12被引 1

对比人类与大模型谈判表现,发现大模型常僵化固守极端立场。

LLM Rationalis? Measuring Bargaining Capabilities of AI Negotiators

  • 用双曲正切函数建模让步行为,提出两个量化指标衡量时机与僵硬度。
  • 大模型在6种权力不对称场景中均锚定在协议区极端,不随形势调整。
  • 模型越大越不会变通,且策略单一,偶有欺骗性行为,适合研究人机博弈者。

双边谈判是复杂且依赖上下文的任务,人类会动态调整锚点、节奏和灵活性以利用权力不对称和非正式线索。本文提出一种基于双曲正切曲线的统一数学框架,用于建模让步动态,并引入两个度量指标:突发性(burstiness tau)和让步刚性指数(Concession-Rigidity Index, CRI),以量化报价轨迹的时间特征与僵硬度。我们在自然语言与数值报价场景下,对人类谈判者与四个前沿大语言模型(LLMs)进行了大规模实证比较,涵盖丰富市场背景及六种受控权力不对称情景。结果表明,与人类能平滑适应情境并推断对手立场与策略不同,大模型系统性地将锚点固定在可能协议区的极端位置,并优化于固定点,无视杠杆或上下文变化。定性分析进一步显示,大模型策略多样性有限,偶有欺骗性行为。此外,模型性能并未随模型规模提升而改善。这些发现揭示了当前大模型谈判能力的根本局限,强调需要更擅长内化对手推理与情境依赖策略的模型。

原文摘要 · Abstract (English)

Bilateral negotiation is a complex, context-sensitive task in which human negotiators dynamically adjust anchors, pacing, and flexibility to exploit power asymmetries and informal cues. We introduce a unified mathematical framework for modeling concession dynamics based on a hyperbolic tangent curve, and propose two metrics burstiness tau and the Concession-Rigidity Index (CRI) to quantify the timing and rigidity of offer trajectories. We conduct a large-scale empirical comparison between human negotiators and four state-of-the-art large language models (LLMs) across natural-language and numeric-offers settings, with and without rich market context, as well as six controlled power-asymmetry scenarios. Our results reveal that, unlike humans who smoothly adapt to situations and infer the opponents position and strategies, LLMs systematically anchor at extremes of the possible agreement zone for negotiations and optimize for fixed points irrespective of leverage or context. Qualitative analysis further shows limited strategy diversity and occasional deceptive tactics used by LLMs. Moreover the ability of LLMs to negotiate does not improve with better models. These findings highlight fundamental limitations in current LLM negotiation capabilities and point to the need for models that better internalize opponent reasoning and context-dependent strategy.

AI谈判大模型评估行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。