arXiv:2601.09032cs.AI2026-01被引 6

评测大模型在真实电商环境中的多步任务能力,发现五层智能层级。

The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments

  • 构建真实电商强化学习环境,测试150项职场任务。
  • 顶级模型仍失败40%,瓶颈在上下文推理而非基础操作。
  • 提出任务导向设计法,适合评估智能体开发与落地。

大语言模型驱动的智能体发展使AI评估从单轮响应转向交互式多步任务完成。我们通过实证研究,在Surge提供的真实电商强化学习环境中,对前沿AI模型在150个职场任务上的表现进行评估。分析揭示了一个经验性得出的“智能体能力层级”:(1) 工具使用,(2) 规划与目标设定,(3) 适应性,(4) 稳定性,(5) 常识推理。即使表现最佳的模型也约有40%的任务失败,且错误集中于该层级结构中。弱模型主要卡在工具使用和规划阶段,而强模型的失败多出现在需要超出明确指令的上下文推断任务上。我们提出以任务为中心的强化学习环境设计方法,强调任务多样性与领域专家参与,提供详细失败分析,并讨论对智能体发展的启示。结果表明,当前前沿模型虽能展现连贯的多步行为,但在真实职场场景中实现人类级任务完成仍存在显著能力缺口。

原文摘要 · Abstract (English)

The advancement of large language model (LLM) based agents has shifted AI evaluation from single-turn response assessment to multi-step task completion in interactive environments. We present an empirical study evaluating frontier AI models on 150 workplace tasks within a realistic e-commerce RL environment from Surge. Our analysis reveals an empirically-derived \emph{hierarchy of agentic capabilities} that models must master for real-world deployment: (1) tool use, (2) planning and goal formation, (3) adaptability, (4) groundedness, and (5) common-sense reasoning. Even the best-performing models fail approximately 40\% of the tasks, with failures clustering predictably along this hierarchy. Weaker models struggle with fundamental tool use and planning, whereas stronger models primarily fail on tasks requiring contextual inference beyond explicit instructions. We introduce a task-centric design methodology for RL environments that emphasizes diversity and domain expert contributions, provide detailed failure analysis, and discuss implications for agent development. Our findings suggest that while current frontier models can demonstrate coherent multi-step behavior, substantial capability gaps remain before achieving human-level task completion in realistic workplace settings.

智能体评估强化学习大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。