arXiv:2512.15431cs.CV2025-12被引 32

用自进化训练流提高GUI自动化数据质量,降低成本90%以上。

Step-GUI Technical Report

  • 通过校准步骤奖励系统,将模型生成轨迹转化为可靠训练信号。
  • 8B模型在AndroidWorld上达80.2%准确率,数据成本降低10-100倍。
  • 提出GUI-MCP协议,支持本地化执行,保障用户隐私安全。

多模态大模型的进展为GUI自动化带来了新机遇,但高质量训练数据的获取仍面临效率与标注可靠性挑战。本文提出一种由校准步骤奖励系统驱动的自进化训练流水线,通过轨迹级校准将模型生成轨迹转化为可靠训练信号,在10-100倍成本降低下实现超90%的标注准确率。基于此,我们构建了Step-GUI系列模型(4B/8B),在主流基准上表现卓越:8B模型在AndroidWorld达80.2%,OSWorld达48.5%,ScreenShot-Pro达62.6%,同时保持强泛化能力。为应对实际部署中跨设备接口统一与隐私保护需求,提出首个面向GUI自动化的模型上下文协议GUI-MCP,采用分层架构融合原子操作与任务委派,实现敏感数据本地处理。为评估真实场景下的使用能力,引入AndroidDaily基准,涵盖3146个静态动作和235个端到端任务,覆盖高频日常场景;8B模型在静态任务上达89.91%,端到端任务达52.50%。本工作推动了实用化GUI代理的发展,展现出在日常数字交互中的广阔应用前景。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models unlock unprecedented opportunities for GUI automation. However, a fundamental challenge remains: how to efficiently acquire high-quality training data while maintaining annotation reliability? We introduce a self-evolving training pipeline powered by the Calibrated Step Reward System, which converts model-generated trajectories into reliable training signals through trajectory-level calibration, achieving >90% annotation accuracy with 10-100x lower cost. Leveraging this pipeline, we introduce Step-GUI, a family of models (4B/8B) that achieves state-of-the-art GUI performance (8B: 80.2% AndroidWorld, 48.5% OSWorld, 62.6% ScreenShot-Pro) while maintaining robust general capabilities. As GUI agent capabilities improve, practical deployment demands standardized interfaces across heterogeneous devices while protecting user privacy. To this end, we propose GUI-MCP, the first Model Context Protocol for GUI automation with hierarchical architecture that combines low-level atomic operations and high-level task delegation to local specialist models, enabling high-privacy execution where sensitive data stays on-device. Finally, to assess whether agents can handle authentic everyday usage, we introduce AndroidDaily, a benchmark grounded in real-world mobile usage patterns with 3146 static actions and 235 end-to-end tasks across high-frequency daily scenarios (8B: static 89.91%, end-to-end 52.50%). Our work advances the development of practical GUI agents and demonstrates strong potential for real-world deployment in everyday digital interactions.

GUI自动化大模型隐私保护基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。