AutoGLM让AI自主操控网页和手机界面,提升智能体在真实场景的交互能力。
AutoGLM: Autonomous Foundation Agents for GUIs
- 设计中间接口分离规划与执行,提升控制灵活性与准确性
- 自进化在线课程强化学习框架,支持持续自我优化
- 在网页和安卓应用中取得超90%成功率,适合实际部署
我们提出AutoGLM,是ChatGLM系列的新成员,旨在作为通过图形用户界面(GUI)自主控制数字设备的基础智能体。尽管基础模型擅长获取人类知识,但在动态现实环境中决策能力不足,制约了通用人工智能的发展。为此,我们聚焦网页浏览器和手机两大典型GUI场景,构建了可落地的AutoGLM基础智能体系统。该系统集成多项技术与基础设施,实现真实世界中可部署的智能体应用。研究得出两大关键洞察:其一,设计合适的“中间接口”对GUI控制至关重要,可分离规划与具身行为,分别优化灵活性与准确性;其二,提出一种新型渐进式训练框架,支持AutoGLM进行自进化在线课程强化学习。评估显示,其在多领域表现优异:网页浏览任务中,于VAB-WebArena-Lite上达55.2%成功率(第二次尝试提升至59.1%),在OpenTable任务中达96.2%;在安卓设备控制中,于AndroidLab(VAB-Mobile)上达36.2%成功率,在主流中文应用常见任务中达89.7%。
原文摘要 · Abstract (English)
We present AutoGLM, a new series in the ChatGLM family, designed to serve as foundation agents for autonomous control of digital devices through Graphical User Interfaces (GUIs). While foundation models excel at acquiring human knowledge, they often struggle with decision-making in dynamic real-world environments, limiting their progress toward artificial general intelligence. This limitation underscores the importance of developing foundation agents capable of learning through autonomous environmental interactions by reinforcing existing models. Focusing on Web Browser and Phone as representative GUI scenarios, we have developed AutoGLM as a practical foundation agent system for real-world GUI interactions. Our approach integrates a comprehensive suite of techniques and infrastructures to create deployable agent systems suitable for user delivery. Through this development, we have derived two key insights: First, the design of an appropriate "intermediate interface" for GUI control is crucial, enabling the separation of planning and grounding behaviors, which require distinct optimization for flexibility and accuracy respectively. Second, we have developed a novel progressive training framework that enables self-evolving online curriculum reinforcement learning for AutoGLM. Our evaluations demonstrate AutoGLM's effectiveness across multiple domains. For web browsing, AutoGLM achieves a 55.2% success rate on VAB-WebArena-Lite (improving to 59.1% with a second attempt) and 96.2% on OpenTable evaluation tasks. In Android device control, AutoGLM attains a 36.2% success rate on AndroidLab (VAB-Mobile) and 89.7% on common tasks in popular Chinese APPs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。