测试大模型代理在动态拍卖中的定价能力,发现顶尖模型表现参差不齐。
Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce

- 构建动态密封投标基准Bazaar,模拟真实市场多属性拍卖场景
- 顶级模型如Gemini 3.1 Pro在客户获取上领先,但利润表现不佳
- 模型对需求突变反应差异大,部分模型恢复速度远超其他
智能体商业正从概念走向实际部署:支付网络、零售商和AI平台正在为代理代表商户与消费者交易铺路。然而,这些代理背后的大型语言模型(LLM)是否能在客户偏好隐蔽、竞争者实时调整、需求突变的现实市场中具备有效定价能力,尚未被系统检验。本文提出Bazaar——一个在动态条件下针对多属性拍卖的密封投标基准。尽管环境动态复杂,该基准基于闭式客户效用函数,可实现精确评估。在来自四家厂商的11个前沿大模型中,领先客户获取的代理(如Gemini 3.1 Pro)未必是利润最高的(如Opus 4.6)。需求冲击下排名再次变化:事前学习最快模型往往事后修正信念最慢,而Gemini 3.1 Pro虽非利润领先者,却恢复最快。然而,即使最强代理也仅获得不到三分之一的后见最优利润,表明当前大模型在智能体商业中虽有进展,但仍存在巨大提升空间。
原文摘要 · Abstract (English)
Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, where customer preferences are hidden, competitors adapt in real time, and demand can shift without warning, has not been systematically tested. We introduce Bazaar, a dynamic sealed-bid benchmark for multi-attribute auction under these conditions. Despite its dynamics, the benchmark is grounded in closed-form customer utilities, enabling exact evaluation. Across 11 frontier LLMs from four providers, the leading agents on customer acquisition (e.g. Gemini 3.1 Pro) are often not the leading agents on profit (e.g. Opus 4.6). The ranking shifts again under demand shocks: agents that learned fastest pre-shock are typically the slowest to revise their beliefs afterwards, while Gemini 3.1 Pro recovers fastest despite not leading on profit. However, even the strongest agent captures less than a third of hindsight-optimal profit, suggesting current LLMs are progressing in agentic commerce but leave substantial headroom.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。