用多模态自动验证提升网页智能体自纠错能力,任务完成率提升5%。
Multimodal Auto Validation For Self-Refinement in Web Agents
- 融合文本与视觉信息自动验证网页操作
- 在WebVoyager基准上任务完成率从76.2%提升至81.24%
- 适合构建高可靠性的自动化数字助手
随着世界数字化进程加快,能够自动化复杂重复任务的网页智能体正成为优化工作流的关键。本文提出一种通过多模态验证与自精炼机制提升网页智能体性能的方法。我们系统研究了文本与视觉模态在自动验证中的作用,以及层级结构的影响,并基于当前最先进的Agent-E网页自动化框架进行改进。引入自精炼机制,使智能体能利用自动验证器检测并修正流程失败。实验表明,该方法在WebVoyager基准的子集上,将Agent-E的任务完成率从76.2%提升至81.24%,显著优于现有水平。本研究为复杂真实场景下更可靠的数字助理提供了新路径。
原文摘要 · Abstract (English)
As our world digitizes, web agents that can automate complex and monotonous tasks are becoming essential in streamlining workflows. This paper introduces an approach to improving web agent performance through multi-modal validation and self-refinement. We present a comprehensive study of different modalities (text, vision) and the effect of hierarchy for the automatic validation of web agents, building upon the state-of-the-art Agent-E web automation framework. We also introduce a self-refinement mechanism for web automation, using the developed auto-validator, that enables web agents to detect and self-correct workflow failures. Our results show significant gains on Agent-E's (a SOTA web agent) prior state-of-art performance, boosting task-completion rates from 76.2\% to 81.24\% on the subset of the WebVoyager benchmark. The approach presented in this paper paves the way for more reliable digital assistants in complex, real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。