评测大模型在真实超市运营中的长期决策能力,发现多数模型难持久稳定
RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments
- 构建千日级超市模拟环境,考验模型跨周期决策能力
- 仅少数模型能撑过180天,最强模型仍远低于理想策略表现
- 适合研究长期自主决策的算法与策略稳定性
大语言模型在短周期任务中进展迅速,但在动态长周期环境中的持续决策能力仍不明确。我们提出RetailBench,一个基于真实数据的单店超市运营仿真基准,用于评估工具调用型大模型代理的长期决策与策略稳定性。该环境将零售管理建模为部分可观测决策过程,支持千日级模拟。代理需应对定价、补货、供应商选择、货架配置、库存老化、客户反馈、外部事件及现金流约束等复杂因素。我们在180天评估期内,对七种主流大模型在典型代理框架下进行测试,并与理想基准策略对比。结果表明模型间差异显著:仅小部分模型能完成全程,即使最强模型在最终净收益和销售表现上也明显落后于基准策略。行为分析揭示差距源于证据获取不全、表面化决策及缺乏一致的长期策略。RetailBench为研究经济驱动下的可靠长期自主决策提供了可控实验平台。
原文摘要 · Abstract (English)
Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain. We introduce RetailBench, a data-grounded simulation benchmark for evaluating tool-using LLM agents in single-store supermarket operation. RetailBench models retail management as a partially observable decision process and is designed to support thousand-day-scale simulations. In this environment, agents must manage pricing, replenishment, supplier selection, shelf assortment, inventory aging, customer feedback, external events, and cash-flow constraints. We evaluate seven contemporary LLMs under representative agent frameworks over a 180-day evaluation horizon and compare them with a privileged oracle policy. Results show substantial variation across models: only a small subset survives the full evaluation horizon, and even the strongest LLM runs remain substantially behind the oracle policy in final net worth and sales outcomes. Behavioral analysis attributes these gaps to incomplete evidence acquisition, surface-level decision making, and the lack of a consistent long-horizon policy. RetailBench provides a controlled testbed for studying reliable autonomy in economically grounded long-horizon decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。