综述大模型驱动的网页智能体,助力自动化日常网络任务
A Survey of WebAgents: Towards Next-Generation AI Agents for Web Automation with Large Foundation Models

- 基于大模型构建可自主执行网页操作的智能体架构
- 现有方法已实现按指令完成复杂网页任务,提升效率
- 适合对AI自动化、人机交互感兴趣的开发者与研究者
随着网页技术的发展,其已深刻改变人们的生活方式。尽管网页至关重要,但许多操作重复且耗时,降低了生活质量。为高效处理这些繁琐任务,基于人工智能的自主智能体(即AI Agent)成为有前景的解决方案,尤其在网页场景中被称为WebAgents,能自动协助用户完成日常任务,显著提升生产力。近年来,包含数十亿参数的大规模基础模型(LFMs)展现出类人语言理解与推理能力,具备完成复杂任务的潜力。这引发关键问题:能否利用LFMs构建强大的自动化网页智能体,为用户提供便利?为此,大量研究聚焦于开发能根据用户指令完成日常网页任务的WebAgents,极大提升了生活便利性。本文系统综述了当前WebAgents在架构、训练与可信性三个方面的研究成果,并探讨未来潜在研究方向,提供深入洞见。
原文摘要 · Abstract (English)
With the advancement of web techniques, they have significantly revolutionized various aspects of people's lives. Despite the importance of the web, many tasks performed on it are repetitive and time-consuming, negatively impacting overall quality of life. To efficiently handle these tedious daily tasks, one of the most promising approaches is to advance autonomous agents based on Artificial Intelligence (AI) techniques, referred to as AI Agents, as they can operate continuously without fatigue or performance degradation. In the context of the web, leveraging AI Agents -- termed WebAgents -- to automatically assist people in handling tedious daily tasks can dramatically enhance productivity and efficiency. Recently, Large Foundation Models (LFMs) containing billions of parameters have exhibited human-like language understanding and reasoning capabilities, showing proficiency in performing various complex tasks. This naturally raises the question: `Can LFMs be utilized to develop powerful AI Agents that automatically handle web tasks, providing significant convenience to users?' To fully explore the potential of LFMs, extensive research has emerged on WebAgents designed to complete daily web tasks according to user instructions, significantly enhancing the convenience of daily human life. In this survey, we comprehensively review existing research studies on WebAgents across three key aspects: architectures, training, and trustworthiness. Additionally, several promising directions for future research are explored to provide deeper insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。