构建大规模数据集与评估工具,系统研究视觉语言模型在手机界面导航中的训练与推理能力。
Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation

- 构建超1.6万任务的HyperTrack数据集和GUIEvalKit评估工具
- 强化学习微调比监督微调更优,尤其在跨领域任务中表现突出
- 揭示交互历史与推理能力对任务完成率的关键影响,适合移动端智能体研究者
视觉语言模型(VLMs)在手机界面导航任务中进展迅速。本文系统研究了该领域的数据规模、基准测试与推理能力。为实现严谨评估,提出HyperTrack——一个包含超过16000个真实世界任务、覆盖650多个中文移动应用的大规模数据集,以及GUIEvalKit——一个开源工具包,支持对离线界面导航任务中VLMs的统一基准测试。基于HyperTrack,分析了训练数据规模对监督与强化学习微调的影响。结果表明,强化学习微调在跨域设置下持续优于监督微调,凸显数据规模与强化学习之间的协同效应。借助GUIEvalKit,进一步对当前最优(SOTA)VLMs进行基准测试,分析交互历史与推理能力对任务完成率的影响。HyperTrack与GUIEvalKit共同构成一个全面的平台,用于开发与评估面向手机界面导航的VLM智能体。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have shown rapid progress in mobile GUI navigation. This paper presents a systematic study of data scaling, benchmarking, and reasoning for VLM-based agents in this domain. To facilitate rigorous evaluation, we introduce HyperTrack, a large-scale dataset with over 16000 real-world tasks across more than 650 Chinese mobile applications, along with GUIEvalKit, an open-source toolkit for unified benchmarking of VLMs on offline GUI navigation tasks. Using HyperTrack, we analyze the effects of training data scale on both supervised and reinforcement-based finetuning. Our results show that reinforcement-based finetuning consistently outperforms supervised finetuning, particularly in out-of-domain settings, highlighting the synergy between data scaling and reinforcement learning. Leveraging GUIEvalKit, we further benchmark state-of-the-art (SOTA) VLMs and analyze how interaction history and reasoning capabilities influence task completion. Together, HyperTrack and GUIEvalKit provide a comprehensive platform for developing and evaluating VLM agents in mobile GUI navigation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。