构建可复现的高保真图形界面环境,助力大模型代理高效训练与评估。
ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis

- 通过大规模合成环境生成技术,实现跨平台GUI交互环境的快速搭建。
- 120个挑战性任务覆盖63个模拟移动应用,平均成功率仅27.92%,远低于人类的92.08%。
- 环境无需后端支持,低资源、零配置,适合大规模训练与评测,适配移动端、桌面端等场景。
基于大语言模型的图形界面代理发展迅速,但真实环境评估面临复杂性高、状态不可控、奖励难定义等问题。现有方法多依赖开源应用或文件操作任务,难以贴近真实使用。虚拟机或Docker方案资源消耗大、响应慢,限制效率。本文提出/sys框架,可生成跨平台、高保真、带可验证奖励的合成交互环境。这些环境以无后端网页形式通过URL访问,几乎零设置且资源开销极低,适用于大规模评估与下游训练。支持移动端、桌面端及车载界面,涵盖100多个环境和1000多个可验证任务。其中120个挑战性任务来自63个模拟移动应用,构成一个完全合成的移动GUI代理基准。五种先进移动端代理测试显示,平均成功率为27.92%,长序列任务下降至17.82%,而人类达92.08%。与真实任务对比表明,合成环境评估结果具有现实泛化能力。项目主页:https://scalewob.github.io。
原文摘要 · Abstract (English)
GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic environments. However, directly doing so in real-world environments introduces some challenges that cannot be overlooked. Real-world environments are complex and uncontrollable, making it difficult to construct verifiable rewards and to save or reset states. Existing works prioritize reproducibility but are often limited to open-source apps or file-operation tasks for reliable reward building, leaving a persistent gap from real-world usage. Furthermore, relying on virtual machines or docker images demand high resource requirements and suffer from slow response speeds, which limit the efficiency. We present \sys, a framework that could produce high-fidelity synthesized interactive environments for GUI agents across platforms with verifiable rewards. These environments behave as backend-free webpages accessible via URL, requiring near-zero setup and low resource cost, making the approach suitable for both large-scale evaluation and downstream agent training. We support multiple GUI platforms including mobile, desktop, and automotive/in-vehicle interfaces based on the same pipeline, covering 100+ environments and 1000+ verifiable tasks. Among them, 120 challenging tasks across 63 simulated mobile applications are released as a fully synthesized mobile GUI agent benchmark. Experiment results on five state-of-the-art mobile GUI agents reveal substantial headroom -- the average success rate is only 27.92\%, dropping to 17.82\% on long-horizon subset -- while humans reach 92.08\%. A comparison against real-world sample tasks shows that assessments made in our synthetic environments generalize to real apps. The project website is at https://scalewob.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。