arXiv:2605.22855cs.GTcs.AI2026-05

测试大模型在隐藏买家偏好下的定价谈判能力,发现能成交但赚不到钱。

PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations

  • 用模拟器构建隐藏偏好谈判环境,限定模型只能输出固定格式动作。
  • 大模型成交率超99%,但利润远低于简单让步策略。
  • 适合研究智能谈判、个性化定价或评估大模型商业决策能力的人。

个性化定价谈判是检验大模型代理能力的难题,因有效互动不等于盈利决策。卖家可能生成合理行为并达成大量交易,却因买家支付意愿和议价特质未知而定价不佳。本文提出PrefBench,一个基于模拟器的隐藏偏好个性化定价谈判基准。每轮对局将模拟买家与固定车辆定制组合配对;卖家可观察公开人格特征、组合信息及谈判历史,但买家的估值、耐心、还价行为和退出决策由隐变量决定。通过面向大模型的状态摘要协议,约束代理在固定隐藏信息边界下返回严格JSON动作。我们在7500轮中评估零样本大模型卖家与启发式基线的表现。测试的大模型可靠遵循协议,成交率高于0.99,但其平均利润仅略高于随机基线,远低于简单让步策略。结果表明,结构化动作合规与达成协议行为可与弱利润敏感的议价能力共存。PrefBench为评估隐藏买家偏好下的定价代理行为提供可控基准。

原文摘要 · Abstract (English)

Personalized pricing negotiations are a challenging testbed for LLM agents because successful interaction does not guarantee profitable decision making. A seller may produce valid actions and close many deals while still pricing poorly when buyer willingness to pay and bargaining traits remain hidden. This paper presents PrefBench, a simulator-based benchmark for hidden-preference personalized pricing negotiations. Each episode pairs a simulated buyer with a fixed vehicle-customization bundle; the seller observes public persona descriptors, bundle information, and negotiation history, while latent buyer variables govern valuation, patience, counter-offer behavior, and walkaway decisions. PrefBench evaluates this setting through an LLM-facing state-summary protocol that constrains agents to return strict JSON actions under a fixed hidden-information boundary. We evaluate zero-shot LLM sellers against heuristic references over 7,500 episodes. The tested LLMs follow the protocol reliably and achieve deal rates above 0.99, but their seller-profit outcomes remain weak: the best LLM average profit is only slightly above the random baseline and far below a simple concession heuristic under the same episode stream. These results show that structured action compliance and agreement-seeking behavior can coexist with weak profit-sensitive bargaining. PrefBench provides a controlled benchmark for evaluating pricing-agent behavior under hidden buyer preferences.

大模型评估谈判系统个性化定价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。