研究自动购物代理如何在有限时间内择机购买,提升消费者收益。
Strategic Buying Agents

- 设计三类策略:静态、贝叶斯和鲁棒,对应不同信息条件下的最优购买阈值。
- 在亚马逊数据上测试,静态与贝叶斯策略平均收益表现良好,鲁棒策略在极端情况更优。
- 语言模型更适合选择策略与校准样本,而非直接做买或等决策。
智能代理正在将线上购物从搜索转向委托购买,即自主代理监控市场并代消费者决定购买时机。本文研究此类策略性购买代理的设计,需在有限购物窗口内决定何时购买,将价格观测、剩余时间与未来价格变化信念转化为购买策略。我们针对三种信息状态——平稳、贝叶斯与鲁棒——建模,将最优策略作为可实施的策略菜单。在平稳状态下,价格调整服从泊松过程且调整后价格分布已知,最优策略为动态阈值规则,阈值由常微分方程决定;在贝叶斯状态下,调整强度已知但调整分布未知,最优规则仍为阈值型,依赖后验信念,并给出知晓真实分布的价值上限;在鲁棒状态下,代理仅知价格上下界,追求最坏情况保护,随机阈值策略可实现最优竞争比与极小最大遗憾。我们在Keepa提供的亚马逊价格历史(367个商品,48,933条时间戳数据)上评估这些策略,并考察其集成至语言模型购买代理的可行性。结果显示,平稳与贝叶斯策略在平均归一化消费者盈余上表现优异,尽管假设简化;而鲁棒策略在分布第10百分位表现最佳。结果表明,语言模型更适合在不同策略与校准样本间进行选择,而非直接做买或等决策。
原文摘要 · Abstract (English)
Agentic AI is shifting online shopping from search toward delegated purchasing, where autonomous buying agents monitor markets and decide when to buy on a consumer's behalf. We study the design of such strategic buying agents, which must decide when to purchase within a finite shopping window, translating price observations, the remaining time horizon, and beliefs about future price changes into a purchase policy. We formulate this problem across three information regimes: stationary, Bayesian, and robust, and treat the resulting optimal policies as a policy menu for implementation. In the stationary regime, price adjustments follow a Poisson arrival process with a known post-adjustment price distribution; the optimal policy is a dynamic purchase-threshold rule, with the threshold governed by an ordinary differential equation. In the Bayesian regime, the adjustment intensity is known, but the price-adjustment distribution is uncertain; the optimal rule remains threshold-based, now depending on posterior beliefs, and we bound the value of knowing the true distribution. In the robust regime, the agent has only price bounds and seeks worst-case protection; randomized threshold policies achieve optimal competitive-ratio and minimax-regret guarantees. We evaluate the proposed policies on Amazon price histories from Keepa (367 items, 48,933 timestamped observations) and examine their integration into language-model buying agents. The stationary and Bayesian policies perform competitively on mean normalized consumer surplus despite their stylized assumptions, while the robust policy performs best at the distribution's 10th percentile. Results suggest language models are better suited to selecting among regimes and calibration samples than to making buy-or-wait decisions directly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。