arXiv:2604.17821cs.AI2026-04ACL被引 2

通过双重不确定性建模,提升网页智能体在复杂任务中的规划与决策能力。

WebUncertainty: Dual-Level Uncertainty Driven Planning and Reasoning For Autonomous Web Agent

论文配图:WebUncertainty: Dual-Level Uncertainty Driven Planning and Reasoning For Autonomous Web Agent
图 1 · 摘自论文原文
  • 基于任务与动作的双重不确定性,动态调整规划与推理策略。
  • 在WebArena和WebVoyager上超越现有方法,显著提升长序列任务成功率。
  • 适合需要高可靠性、自主执行复杂网页操作的研究者与开发者。

大型语言模型(LLMs)的进步使自主网页智能体能够直接在真实网页上执行自然语言指令。然而,现有智能体在涉及动态交互和长时程执行的复杂任务中表现不佳,主要受限于僵化的规划策略和易产生幻觉的推理过程。为此,我们提出WebUncertainty,一种针对规划与推理中双重不确定性设计的新框架。具体而言,设计了任务不确定性驱动的自适应规划机制,可依据环境未知性动态选择规划模式;同时引入动作不确定性驱动的蒙特卡洛树搜索(MCTS)推理机制,结合置信度驱动的动作不确定性(ConActU)策略,量化了偶然性不确定性(AU)与认知性不确定性(EU),从而优化搜索路径并指导鲁棒决策。在WebArena与WebVoyager基准上的实验表明,WebUncertainty在性能上优于当前最先进基线。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have empowered autonomous web agents to execute natural language instructions directly on real-world webpages. However, existing agents often struggle with complex tasks involving dynamic interactions and long-horizon execution due to rigid planning strategies and hallucination-prone reasoning. To address these limitations, we propose WebUncertainty, a novel autonomous agent framework designed to tackle dual-level uncertainty in planning and reasoning. Specifically, we design a Task Uncertainty-Driven Adaptive Planning Mechanism that adaptively selects planning modes to navigate unknown environments. Furthermore, we introduce an Action Uncertainty-Driven Monte Carlo tree search (MCTS) Reasoning Mechanism. This mechanism incorporates the Confidence-induced Action Uncertainty (ConActU) strategy to quantify both aleatoric uncertainty (AU) and epistemic uncertainty (EU), thereby optimizing the search process and guiding robust decision-making. Experimental results on the WebArena and WebVoyager benchmarks demonstrate that WebUncertainty achieves superior performance compared to state-of-the-art baselines.

自主智能体不确定性建模网页操作强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。