提出一步校准框架,提升网页导航准确率。
StepGuard: Guarding Web Navigation via Single-Step Calibration

- 双策略动态切换,分离探索与答题任务
- 按置信度触发反思,用对比奖励自纠正
- 在标准数据集上达到新最优表现
网页导航需代理遵循自然语言目标,与网页交互并生成准确答案。尽管近期进展结合了视觉-语言模型与强化学习,现有方法仍因奖励错配和误差传播导致单步脆弱性。为缓解奖励纠缠,我们设计动态双策略优化(DDPO),动态切换探索优先与答案优先模式以减少奖励冲突。为校准单步误差,提出置信度引导的自适应导航反思机制(CANR),通过每步置信度估计,仅在必要时触发反思,并利用对比奖励促进自我修正以校准单步偏差。基于上述组件,我们构建了StepGuard框架,实现网页导航的单步校准。实验表明,该方法显著提升导航与答案准确率,在标准网页导航基准上达到新的最优性能。
原文摘要 · Abstract (English)
Web navigation requires agents to follow natural language goals, interact with web pages, and produce accurate answers. While recent advances leverage vision-language models and reinforcement learning, existing methods still suffer from single-step fragility due to reward misalignment and error propagation. To tackle the reward entanglement, we design Dynamic Dual-Policy Optimization (DDPO), which dynamically switches between a navigation-first mode for exploration and an answer-first mode for question-answering to mitigate reward conflict. To calibrate the single-step error, we propose Confidence-Guided Adaptive Navigation Reflection (CANR), a mechanism that estimates per-step confidence, triggers reflection only when necessary, and uses contrastive rewards to encourage self-correction to calibrate the single-step inaccuracy. With the above as the main components, we finally develop our StepGuard, a new framework of Guarding Web Navigation via Single-Step Calibration. Experiments demonstrate that our approach significantly improves navigation and answer accuracy, setting new state-of-the-art performance on standard web navigation benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。