arXiv:2606.14397cs.LG2026-06

新基准测试发现顶尖智能体在复杂任务中表现远低于人类。

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

论文配图:Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
图 1 · 摘自论文原文
  • 构建网页版基准,聚焦时间感知、图像理解与3D推理能力。
  • 顶尖智能体仅19.1%成功率,远低于人类超80%水平。
  • 适合评估智能体在真实专业场景下的泛化能力。

随着智能体系统在现实场景中的广泛应用,对其能力进行真实评估的需求日益迫切。然而,现有基准多基于常见应用,任务简单且覆盖能力有限,导致现代智能体性能饱和,难以揭示其真实短板。为此,我们提出GauntletBench,一个面向挑战性场景的网页基准,聚焦时间感知、图形理解与3D推理三大被忽视能力,涵盖视频编辑器、工作流构建器、3D建模器、飞行分析仪和电路设计等五类专业应用,每类含20个视觉密集型任务(共100个)。该基准提供兼容开源与闭源框架的模块化流程,包括适配环境、受控网页应用、结构化任务集及支持多种指标的自动化评估引擎。实证结果表明,前沿智能体在该基准上仅达19.1%成功率达,远未接近人类水平。非专家人类标注者在这些具有挑战性但可行的任务中成功率超80%,凸显当前智能体与真实世界复杂需求间的巨大差距。

原文摘要 · Abstract (English)

As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular applications with relatively simple tasks and focus on a narrow set of capabilities while overlooking broader dimensions, resulting in saturated performance on modern agents and failing to probe their limitations. To this end, we introduce GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), across five less-covered professional applications (Video Editor, Workflow Builder, 3D Modeller, Flight Analyser, and Circuit Designer), each with 20 vision-intensive tasks (100 in total). Our benchmark provides a modular pipeline that comprises an environment compatible with both open- and closed-source agent frameworks, a controlled web-based application, a well-structured task suite, and an automated evaluation engine with diverse metrics. Contrary to widespread expectations, our empirical results reveal that frontier agentic systems remain far from achieving human-level performance. Even the state-of-the-art agent achieves only a 19.1% success rate on our GauntletBench, highlighting the limitations in these overlooked capabilities and generalisation. By comparison, non-expert human annotators achieve over 80% success on our challenging yet feasible tasks, revealing the substantial gap between current agent capabilities and those required for complex real-world scenarios.

智能体评估泛化能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。