arXiv:2603.15039cs.CV2026-03中稿 · CVPR被引 3

首个面向中文手机GUI代理的综合性评估基准,覆盖201款主流应用。

GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents

  • 基于真实设备构建双层评估框架,涵盖感知到执行全链路能力。
  • 在201个主流应用上测试,发现多数模型在反思与自评环节表现不足。
  • 专为中文生态设计,适合研究移动智能体与多模态大模型的开发者。

多模态大语言模型(MLLMs)的发展使具备视觉感知、跨模态推理和交互控制能力的手机GUI代理成为可能。然而,现有基准大多以英语为主,未能反映中文移动生态的语言与交互特性,且仅关注孤立技能如GUI定位或离线代理,缺乏统一、细粒度的全能力链评估框架。为此,我们提出GUI-CEval,首个完全基于真实设备环境的中文手机GUI代理综合评测基准。该基准覆盖4种设备类型下的201款主流应用,采用两级结构,从感知、规划、反思、执行到评估五个维度,全面评估原子能力与实际应用性能。所有数据通过多阶段人工采集与验证,确保真实性与可复现性。对20个代表性MLLMs及多智能体系统的广泛实验表明,尽管Qwen2.5-VL和UI-TARS表现较优,但多数模型在反思决策与动作后自我评估方面仍存在明显短板,限制其在真实交互中的可靠性。我们期望GUI-CEval能为能力诊断提供全面、可解释的评估工具,推动中文手机GUI代理的发展。

原文摘要 · Abstract (English)

Recent progress in Multimodal Large Language Models (MLLMs) has enabled mobile GUI agents capable of visual perception, cross-modal reasoning, and interactive control. However, existing benchmarks are largely English-centric and fail to capture the linguistic and interaction characteristics of the Chinese mobile ecosystem. They also focus on isolated skills such as GUI grounding or offline agent, lacking a unified and fine-grained framework to assess the full capability chain from perception to execution. To address this gap, we introduce GUI-CEval, the first comprehensive benchmark for Chinese mobile GUI agents, built entirely on physical device environments. GUI-CEval spans 201 mainstream apps across four device types and adopts a two-level structure that evaluates both atomic abilities and realistic application-level performance along five dimensions: perception, planning, reflection, execution, and evaluation. All data are collected and verified through multi-stage manual processes to ensure authenticity and reproducibility. Extensive experiments on 20 representative MLLMs and multi-agent systems show that while models such as Qwen2.5-VL and UI-TARS perform competitively, most MLLMs still exhibit clear weaknesses in reflective decision-making and post-action self-evaluation, limiting their reliability in real-world interactions. We hope GUI-CEval provides a comprehensive and interpretable benchmark to guide capability diagnosis and advance the development of Chinese mobile GUI agents.

GUI代理多模态模型中文评估移动智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。