用一年真实电商运营测试大模型,发现没有哪个模型全面领先。
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

- 设计一年周期的电商模拟环境,含多方谈判与突发事件
- GPT-5.6 最终资产达 143万,但防欺诈能力仅第16
- 开源代码,适合评估大模型长期决策与适应能力
长时序智能体任务超越简单任务串联,需在动态环境中持续探索、学习并调整策略。我们提出 E-Commerce Bench,首个开源基准,将多轮对手谈判与动态事件融入为期一年的自主商业运营。在365天内,LLM代理同时管理多个在线店铺,执行市场调研、供应商议价采购、销售策略优化、订单履约、退货处理及现金流管理,以最大化年末总资产。为构建真实商户环境,产品与供应商数据来自真实电商平台,而促销、自然灾害和供应链冲击等年度日历事件持续影响需求。为保证可复现性,市场双方均为确定性:客户购买与退货遵循固定需求模型,谈判内核决定供应商定价、让步与决策,仅用LLM生成语言表达。我们在七项维度上评估18个前沿模型,发现无一模型全面领先。GPT-5.6 收益最高,初始10万资金增长至1,431,425,但欺诈规避排名仅第16,运营效率低于Fable5。开源权重模型中,Qwen3.8-Max-Preview 表现最佳,达416,252,较GLM 5.2(高)高出38%,且随时间逐步降低采购价格,展现出强学习能力。代码已开源:https://github.com/QwenLM/E-CommerceBench。
原文摘要 · Abstract (English)
Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environments and long-range dependencies require Large Language Models (LLMs) to continually explore, learn from experience, and adapt their policies over thousands of steps. We introduce E-Commerce Bench, the first open-source benchmark that integrates multi-round counterpart negotiation and dynamic events into a year-long business operation. Over a 365-day year, an LLM agent concurrently runs multiple online stores, researching the market, negotiating with suppliers to source inventory, optimizing sales strategies, fulfilling orders, handling returns, and managing cash flow to maximize its end-of-year total assets. To construct a realistic merchant-side operating environment, the product and supplier data are derived from a real e-commerce platform, while a year-long calendar of promotions, natural disasters, and supply-chain shocks continually reshapes demand. For reproducibility, both sides of the market are deterministic: customer purchases and returns follow a fixed demand model, while a negotiation kernel determines supplier pricing, concessions, and decisions, with an LLM used only to verbalize them. We evaluate 18 frontier models across seven dimensions, including year-end assets, and find that no single model dominates. GPT-5.6 Sol earns the most, growing the 100,000 opening stake into 1,431,425, yet it ranks 16th of 18 on fraud avoidance and trails Fable5 in operational efficiency. Among open-weight models, Qwen3.8-Max-Preview leads with 416,252, 38% above GLM 5.2 (high), and achieves the strongest learning over the horizon, progressively bargaining down prices across repeated orders. Our code is available at https://github.com/QwenLM/E-CommerceBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。