仅需一次演示即可稳定完成复杂界面操作,速度快10倍且更可靠。
GPA: Learning GUI Process Automation from Demonstrations

- 基于蒙特卡洛定位的序列化方法,抗界面缩放与识别误差。
- 通过就绪状态校准确保执行确定性,成功率显著提升。
- 全程本地运行,适合企业级安全敏感任务,也可作为工具被其他智能体调用。
GUI流程自动化(GPA)是一种轻量级但通用的视觉驱动型机器人流程自动化技术,仅需一次示范即可实现快速稳定的流程重演。针对传统RPA脆弱和现有视觉语言模型类GUI代理的非确定性风险,GPA提出三大核心优势:(1) 采用基于序贯蒙特卡洛的定位方法,有效应对界面缩放与检测不确定性;(2) 通过就绪状态校准保障执行的确定性与可靠性;(3) 实现快速、完全本地化的执行,保护隐私。该方法满足企业工作流所需的适应性、鲁棒性与安全性要求。此外,GPA可作为MCP/CLI工具供具备编码能力的智能体使用,使智能体仅负责逻辑推理与编排,而由GPA处理具体界面操作。我们开展初步实验对比GPA与Gemini 3 Pro(搭配CUA工具),结果表明,在完成长时程GUI任务时,GPA成功率达更高,执行速度提升10倍。
原文摘要 · Abstract (English)
GUI Process Automation (GPA) is a lightweight but general vision-based Robotic Process Automation (RPA), which enables fast and stable process replay with only a single demo. Addressing the fragility of traditional RPA and the non-deterministic risks of current vision language model-based GUI agents, GPA introduces three core benefits: (1) Robustness via Sequential Monte Carlo-based localization to handle rescaling and detection uncertainty; (2) Deterministic and Reliability safeguarded by readiness calibration; and (3) Privacy through fast, fully local execution. This approach delivers the adaptability, robustness, and security required for enterprise workflows. It can also be used as an MCP/CLI tool by other agents with coding capabilities so that the agent only reasons and orchestrates while GPA handles the GUI execution. We conducted a pilot experiment to compare GPA with Gemini 3 Pro (with CUA tools) and found that GPA achieves higher success rate with 10 times faster execution speed in finishing long-horizon GUI tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。