从分层规划视角揭示大模型网页代理失败根源
Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective

- 构建三层分析框架:高层规划、底层执行与重规划
- 自然语言计划比PDDL计划更冗长,底层执行是主要瓶颈
- 提升感知对齐与自适应控制比优化推理更重要
大语言模型(LLM)网页代理在真实长周期任务中仍远未达到人类可靠性。现有评估多聚焦端到端成功率,难以揭示失败原因。本文提出分层规划框架,从高层规划、底层执行和重规划三个层面分析代理行为,实现对推理、对齐和恢复能力的过程化评估。实验表明,结构化的规划域定义语言(PDDL)计划比自然语言(NL)计划更简洁、目标导向更强;但底层执行仍是主要瓶颈。结果表明,提升感知对齐与自适应控制能力,而非仅优化高层推理,对实现人类级可靠性至关重要。该分层视角为诊断和推进LLM网页代理提供了系统性基础。
原文摘要 · Abstract (English)
Large language model (LLM) web agents are increasingly used for web navigation but remain far from human reliability on realistic, long-horizon tasks. Existing evaluations focus primarily on end-to-end success, offering limited insight into where failures arise. We propose a hierarchical planning framework to analyze web agents across three layers (i.e., high-level planning, low-level execution, and replanning), enabling process-based evaluation of reasoning, grounding, and recovery. Our experiments show that structured Planning Domain Definition Language (PDDL) plans produce more concise and goal-directed strategies than natural language (NL) plans, but low-level execution remains the dominant bottleneck. These results indicate that improving perceptual grounding and adaptive control, not only high-level reasoning, is critical for achieving human-level reliability. This hierarchical perspective provides a principled foundation for diagnosing and advancing LLM web agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。