拆解网页智能体瓶颈,发现规划能力是主要短板。
From Grounding to Planning: Benchmarking Bottlenecks in Web Agents
- 分离规划与定位组件,分别构建评估基准
- 实测表明定位非瓶颈,当前技术已可解决
- 规划能力不足才是性能差主因,适合研究者参考
通用网页智能体在复杂网络环境中的交互日益重要,但其在真实应用中表现仍差,即使使用最先进模型,准确率也极低。我们发现这些智能体可分解为规划与定位两大组件,而现有研究多将其视为黑箱,仅做端到端评估,阻碍了有效改进。本文通过在Mind2Web数据集上细化实验,明确区分两个组件,并分别为其设计新基准,识别出制约智能体性能的瓶颈。与普遍认知相反,我们的结果表明定位并非显著瓶颈,现有技术已能有效应对;真正核心问题是规划组件,它是导致性能下降的主要原因。本研究提供新洞见并给出实用改进建议,为打造更可靠的网页智能体铺平道路。
原文摘要 · Abstract (English)
General web-based agents are increasingly essential for interacting with complex web environments, yet their performance in real-world web applications remains poor, yielding extremely low accuracy even with state-of-the-art frontier models. We observe that these agents can be decomposed into two primary components: Planning and Grounding. Yet, most existing research treats these agents as black boxes, focusing on end-to-end evaluations which hinder meaningful improvements. We sharpen the distinction between the planning and grounding components and conduct a novel analysis by refining experiments on the Mind2Web dataset. Our work proposes a new benchmark for each of the components separately, identifying the bottlenecks and pain points that limit agent performance. Contrary to prevalent assumptions, our findings suggest that grounding is not a significant bottleneck and can be effectively addressed with current techniques. Instead, the primary challenge lies in the planning component, which is the main source of performance degradation. Through this analysis, we offer new insights and demonstrate practical suggestions for improving the capabilities of web agents, paving the way for more reliable agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。