arXiv:2609.04667cs.AI2026-09

测试大模型在不同竞争市场中的企业决策能力,发现表现排名会随环境变化。

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

论文配图:ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
图 1 · 摘自论文原文
  • 构建六轮企业资源规划模拟,含定价、生产等环节的动态竞争环境。
  • 同一模型在单人对战和多人竞争中排名差异显著,领先者互换。
  • 适合评估企业AI代理在真实市场竞争中的适应性与鲁棒性。

大型语言模型(LLM)代理被越来越多地应用于企业流程,但现有评估很少检验其商业决策结论是否能在不同竞争市场生态中迁移。我们提出ERPBench,一个执行可追踪的企业决策代理基准,包含六轮企业资源规划(ERP)模拟,涵盖定价、生产、采购、库存、财务及共享市场竞争。该基准在两个匹配的市场生态中评估相同的100个固定问题:Solo模式下,每个待测LLM代理与固定规则对手竞争;Arena模式下,六个待测LLM代理在共享市场中相互竞争。覆盖六种模型家族,共生成1,200条模型级轨迹,总计7,200次决策回合。在观察的服务配置下,领先模型因生态而异:DeepSeek在Solo中表现最优(平均估值252.29M,平均排名1.67),Gemini在Arena中领先(263.95M,1.76)。仅21/100问题在两种生态中识别出相同任务级优胜者,且Gemini在Arena中末位率从22%降至0%。ERPBench支持配对评估企业代理排名在竞争生态间的可迁移性,并提供聚合执行干预分析。代码与数据集见https://github.com/GAIR-NLP/erp-bench。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations rarely test whether business-decision conclusions transfer across competitive market ecologies. We introduce ERPBench, an execution-instrumented benchmark for enterprise decision agents in a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition. ERPBench evaluates the same 100 fixed problems in two matched competitive market ecologies: Solo, where each evaluated LLM agent competes against fixed rule-based opponents, and Arena, where six evaluated LLM agents compete in a shared market. Across six model families, this yields 1,200 model-level trajectories spanning 7,200 decision rounds. Under the observed service configuration, the leading model differs between ecologies: DeepSeek leads in Solo (252.29M mean valuation; mean rank 1.67), whereas Gemini leads in Arena (263.95M; 1.76). The two ecologies identify the same task-level winner on only 21 of 100 problems, and Gemini's bottom-rank rate falls from 22 % to 0 % in Arena. ERPBench supports paired evaluation of whether enterprise-agent rankings transfer across competitive market ecologies, supplemented by aggregate execution-intervention analysis. Code and benchmark resources are available in our https://github.com/GAIR-NLP/erp-bench.

大模型评估企业决策竞争模拟多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。