构建真实跨境电商环境,评估大模型代理的长期经营能力。
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

- 在基于真实阿里数据的模拟市场中运行店铺,考验长期决策与适应力。
- 15个模型间净财富相差九倍,最优仍逊于人类策略。
- 通过技能与动作溯源分析,揭示模型的经营风格与关键失误点。
运营企业是极具挑战的智能任务:需从不完整信息中发现机会,在不确定性下投入资本,适应延迟反馈的动态市场,并满足合规要求才能合法交易。前沿大模型虽能完成复杂流程,但其商业能力鲜有被系统评估。我们提出「Business Arena」——一个受控环境,让AI代理在长周期内运营跨境店铺,从供应商采购并销售给买家。该环境基于真实的阿里巴巴采购数据,市场条件由权威来源校准。由于后果延迟且相互关联,单个决策难以评判,但整体表现可通过利润衡量。为理解成败原因,我们对比人类设计策略以估算潜在机会,使用技能水平指标揭示模型优劣,并追溯收益与损失到具体行为。通过机制消融实验,确认优秀表现反映真实商业智能而非疏忽或仿真捷径。评估15个前沿模型发现,平均最终净资产相差九倍。即使最佳模型也落后于人类策略,表明商业运营对大模型仍具挑战。技能分析揭示不同经营风格,如高利润率精品卖家、高周转批发商和客服专精者;动作层面归因则识别出影响价值的关键采购、定价与补货决策。Business Arena为评估端到端商业代理迈出第一步。
原文摘要 · Abstract (English)
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce \textbf{Business Arena}, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。