GUI智能体需从能力提升转向工程化建设,才能真正落地。
Software Engineering for and with GUI Agent

- 构建模块化感知-推理-执行闭环,增强系统自洽性。
- 当前评估仍以任务成功率为主,缺乏可比性和长期维护支持。
- 适合关注AI系统可部署性、安全与人机协同的研究者。
GUI智能体发展迅速,催生了大量框架、基准和应用,但其技术仍脆弱、工程不完善,缺乏持续真实场景使用的验证。它们正演变为闭环软件系统,模型推理与界面感知、执行反馈、恢复机制及人工监督耦合。这一演变呼唤尚未被充分重视的软件工程视角。本文综述2018年1月至2026年4月间336篇相关论文,围绕研究现状、架构、评估、生命周期问题与未来机遇提出五个研究问题。结果显示,自2024年起领域扩张显著,移动端与网页端仍是主流。架构普遍采用感知-推理-执行的模块化循环,但恢复、人工升级、安全控制与可审计性仍不足。评估也呈现交互性增强趋势,但仍集中于任务成功率,跨协议比较困难。现有研究对基准之外的测试支持有限,发布后维护能力薄弱。可观测性、隐私工程与系统化人工监督同样欠缺。这些发现表明,仅提升能力无法保证部署就绪。未来研究应将可靠执行与生命周期测试、可复现评估相连接,并整合权限控制、隐私保护与成本敏感的人本治理,以构建可信赖、可持续、安全且可部署的GUI智能体系统。
原文摘要 · Abstract (English)
GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity of the field. GUI agents remain technically brittle, incompletely engineered, and insufficiently validated for sustained real-world use. They are evolving into closed-loop software systems. Within these systems, model reasoning is coupled with interface perception, execution feedback, recovery, and human oversight. This evolution calls for a software engineering perspective that remains largely absent from existing research. We address this gap by reviewing 336 GUI-agent papers from January 2018 to April 2026. Five research questions examine the research landscape, architectures, evaluation, software lifecycle concerns, and future opportunities. Our findings show that the field has expanded sharply since 2024, while mobile and web settings remain dominant. Architectures increasingly adopt modular perceive-reason-act loops, but recovery, human escalation, safety enforcement, and auditability remain underdeveloped. This architectural imbalance extends to evaluation. Evaluations are becoming more interactive, but they remain centered on task success and are difficult to compare across protocols. More broadly, existing studies provide limited support for testing beyond benchmarks and for maintaining agents after release. Observability, privacy engineering, and systematic human oversight are also underdeveloped. Together, these findings show that capability improvements alone cannot ensure deployment readiness. Future research should connect dependable execution with lifecycle-centered testing and reproducible evaluation. It should also integrate permission and privacy controls with cost-aware, human-centered governance. This integration is necessary to build dependable, maintainable, secure, and deployable GUI-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。