测试大模型在超市运营中长期决策能力,发现多数模型难以持续表现。
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
- 构建千日级超市模拟环境,评估模型跨周期决策能力
- 仅少数模型能撑过180天,最强模型仍远低于理想策略
- 适合研究长期自主智能、经济决策与代理系统可靠性
大语言模型代理在短期任务上进展迅速,但在动态的长期环境中维持连贯决策的能力仍不明确。我们提出RetailBench,一个基于真实零售场景的数据驱动仿真基准,用于评估工具使用型大模型代理在单店超市运营中的表现。该环境将零售管理建模为部分可观测的决策过程,支持千日级模拟。代理需处理定价、补货、供应商选择、货架组合、库存老化、客户反馈、外部事件及现金流约束。我们在180天评估周期内对七种主流大模型在典型代理框架下进行测试,并与理想化的最优策略(oracle policy)对比。结果表明模型间表现差异显著:仅有少量模型能完成全程评估,即使最强模型在最终净收益和销售额上仍明显落后于理想策略。行为分析显示差距源于证据获取不全、表面决策及缺乏一致的长期策略。RetailBench为研究经济驱动下的长期自主决策提供了可控实验平台。
原文摘要 · Abstract (English)
Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain. We introduce RetailBench, a data-grounded simulation benchmark for evaluating tool-using LLM agents in single-store supermarket operation. RetailBench models retail management as a partially observable decision process and is designed to support thousand-day-scale simulations. In this environment, agents must manage pricing, replenishment, supplier selection, shelf assortment, inventory aging, customer feedback, external events, and cash-flow constraints. We evaluate seven contemporary LLMs under representative agent frameworks over a 180-day evaluation horizon and compare them with a privileged oracle policy. Results show substantial variation across models: only a small subset survives the full evaluation horizon, and even the strongest LLM runs remain substantially behind the oracle policy in final net worth and sales outcomes. Behavioral analysis attributes these gaps to incomplete evidence acquisition, surface-level decision making, and the lack of a consistent long-horizon policy. RetailBench provides a controlled testbed for studying reliable autonomy in economically grounded long-horizon decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。