arXiv:2509.14382cs.AI2025-09被引 1

通过细粒度分析发现网页智能体执行中的隐藏错误

Detecting Pipeline Failures through Fine-Grained Analysis of Web Agents

  • 将智能体流程拆解为可解释阶段,逐环节诊断问题
  • 在Mind2Web数据集上揭示传统指标忽略的执行缺陷
  • 适合研究智能体鲁棒性与系统优化的开发者

由大语言模型驱动的网页智能体可在动态网络环境中自主完成复杂多步任务。然而,当前评估主要关注整体成功率,忽视中间环节的错误,限制了对失败模式的理解,也阻碍了系统的持续改进。本文分析现有基准,指出缺乏细粒度诊断工具的问题。为此,我们提出一个模块化评估框架,将智能体流程分解为可解释的阶段,实现细致的错误分析。以SeeAct框架和Mind2Web数据集为例,验证该方法能揭示标准指标未捕捉到的可操作性弱点,为构建更鲁棒、泛化能力更强的网页智能体铺平道路。

原文摘要 · Abstract (English)

Web agents powered by large language models (LLMs) can autonomously perform complex, multistep tasks in dynamic web environments. However, current evaluations mostly focus on the overall success while overlooking intermediate errors. This limits insight into failure modes and hinders systematic improvement. This work analyzes existing benchmarks and highlights the lack of fine-grained diagnostic tools. To address this gap, we propose a modular evaluation framework that decomposes agent pipelines into interpretable stages for detailed error analysis. Using the SeeAct framework and the Mind2Web dataset as a case study, we show how this approach reveals actionable weaknesses missed by standard metrics - paving the way for more robust and generalizable web agents.

智能体评测框架细粒度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。