arXiv:2601.20335cs.CLcs.AI2026-01ACL被引 4

首个面向中文手机界面智能体的综合评估基准,覆盖真实环境噪声与复杂推理。

MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment

  • 构建1080个来自80款中文App的真实任务,分5类维度评估执行、推理与抗噪能力。
  • 引入自动重置框架,实现稳定可重复的真机环境评测,12个主流智能体表现均不理想。
  • 适合研究移动端智能体、人机交互及真实场景鲁棒性评估的团队使用。

近年来,移动图形用户界面(GUI)智能体的发展凸显了全面评估基准的迫切需求。尽管新出现的在线基准比离线基准更贴近真实场景,但普遍侧重于任务指令遵循能力,忽视了智能体的推理与探索能力,且未考虑真实移动环境中存在的随机噪声,导致评测与实际应用存在差距。为此,我们提出MobileBench-OL,一个包含80款中文应用中1080个任务的在线基准。该基准通过5个子集涵盖任务执行、复杂推理和噪声鲁棒性等多个评估维度。同时,我们设计了一个带重置机制的自动评估框架,确保评测在真实设备上的稳定性与可重复性。在MobileBench-OL上对12个领先GUI智能体的评估显示,其性能仍有显著提升空间以满足真实应用需求。人工评估进一步验证了该基准能可靠衡量主流智能体在真实环境中的表现。数据与代码将在论文录用后公开。

原文摘要 · Abstract (English)

Recent advances in mobile Graphical User Interface (GUI) agents highlight the growing need for comprehensive evaluation benchmarks. While new online benchmarks offer more realistic testing than offline ones, they tend to focus on the agents' task instruction-following ability while neglecting their reasoning and exploration ability. Moreover, these benchmarks do not consider the random noise in real-world mobile environments. This leads to a gap between benchmarks and real-world environments. To addressing these limitations, we propose MobileBench-OL, an online benchmark with 1080 tasks from 80 Chinese apps. It measures task execution, complex reasoning, and noise robustness of agents by including 5 subsets, which set multiple evaluation dimensions. We also provide an auto-eval framework with a reset mechanism, enabling stable and repeatable real-world benchmarking. Evaluating 12 leading GUI agents on MobileBench-OL shows significant room for improvement to meet real-world requirements. Human evaluation further confirms that MobileBench-OL can reliably measure the performance of leading GUI agents in real environments. Our data and code will be released upon acceptance.

移动智能体评估基准真实场景中文应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。