打造能在手机上高效执行图形操作的中文智能代理,性能领先。
AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning
- 用强化学习微调提升模型在界面中的推理能力。
- 在中英文界面测试中达96.9%类型匹配、91.3%精确匹配。
- 专为移动端设计,支持低延迟执行,适合中文应用生态。
大型语言模型代理的进展为通过图形用户界面(GUI)自动化任务带来了新可能,尤其在移动环境中,智能交互能显著提升易用性。然而,实际部署仍受多重挑战制约:现有训练数据噪声多、语义多样性不足,影响精准定位与规划;纯模仿学习易过拟合于已见界面模式,在陌生场景下泛化能力差;多数研究聚焦英文界面,忽视中文等非英语应用生态的多样性。本文提出AgentCPM-GUI,一个80亿参数的GUI代理,专为设备端稳健高效的界面交互设计。训练流程包括感知增强的接地预训练、高质量中英文轨迹的监督微调,以及基于GRPO的强化微调以提升推理能力。同时引入紧凑动作空间,减少输出长度,支持移动端低延迟执行。该模型在五个公开基准及新提出的中文GUI基准CAGUI上表现最优,达到96.9%类型匹配和91.3%精确匹配。为促进复现与研究,所有代码、模型权重与评估数据均已开源。
原文摘要 · Abstract (English)
The recent progress of large language model agents has opened new possibilities for automating tasks through graphical user interfaces (GUIs), especially in mobile environments where intelligent interaction can greatly enhance usability. However, practical deployment of such agents remains constrained by several key challenges. Existing training data is often noisy and lack semantic diversity, which hinders the learning of precise grounding and planning. Models trained purely by imitation tend to overfit to seen interface patterns and fail to generalize in unfamiliar scenarios. Moreover, most prior work focuses on English interfaces while overlooks the growing diversity of non-English applications such as those in the Chinese mobile ecosystem. In this work, we present AgentCPM-GUI, an 8B-parameter GUI agent built for robust and efficient on-device GUI interaction. Our training pipeline includes grounding-aware pre-training to enhance perception, supervised fine-tuning on high-quality Chinese and English trajectories to imitate human-like actions, and reinforcement fine-tuning with GRPO to improve reasoning capability. We also introduce a compact action space that reduces output length and supports low-latency execution on mobile devices. AgentCPM-GUI achieves state-of-the-art performance on five public benchmarks and a new Chinese GUI benchmark called CAGUI, reaching $96.9\%$ Type-Match and $91.3\%$ Exact-Match. To facilitate reproducibility and further research, we publicly release all code, model checkpoint, and evaluation data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。