arXiv:2512.12634cs.AI2025-12被引 3

MobiBench让手机界面智能体评估更真实、可复现,支持多路径和模块化分析。

MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents

  • 构建多路径、模块化离线评测框架,避免单路径偏见
  • 与人类评估者一致性达94.72%,媲美在线评测但更可复现
  • 揭示模型组件瓶颈,提供高效设计指南,适合开发者优化智能体

移动GUI代理具备代表用户操作手机应用的潜力,但现有评估方法存在两大缺陷:一是依赖单一路径的离线或动态在线评测,前者因固定路径惩罚合理替代动作,后者因环境不可预测导致可扩展性差;二是将代理视为黑箱,忽略各组件贡献,造成评价不公或性能瓶颈隐藏。为此,我们提出MobiBench,首个面向移动GUI代理的多路径、模块化离线评测框架,实现高保真、可扩展、可复现的全离线评估。实验表明,MobiBench与人工评估者达成94.72%的一致性,媲美精心设计的在线评测,同时保留静态离线评测的可复现优势。进一步模块级分析揭示了多种技术的有效性、不同模型规模下的最优组件配置、当前大型语言模型的固有局限,并提供了设计更高效智能体的实用建议。

原文摘要 · Abstract (English)

Mobile GUI Agents, AI agents capable of interacting with mobile applications on behalf of users, have the potential to transform human computer interaction. However, current evaluation practices for GUI agents face two fundamental limitations. First, they either rely on single path offline benchmarks or online live benchmarks. Offline benchmarks using static, single path annotated datasets unfairly penalize valid alternative actions, while online benchmarks suffer from poor scalability and reproducibility due to the dynamic and unpredictable nature of live evaluation. Second, existing benchmarks treat agents as monolithic black boxes, overlooking the contributions of individual components, which often leads to unfair comparisons or obscures key performance bottlenecks. To address these limitations, we present MobiBench, the first modular and multi path aware offline benchmarking framework for mobile GUI agents that enables high fidelity, scalable, and reproducible evaluation entirely in offline settings. Our experiments demonstrate that MobiBench achieves 94.72 percent agreement with human evaluators, on par with carefully engineered online benchmarks, while preserving the scalability and reproducibility of static offline benchmarks. Furthermore, our comprehensive module level analysis uncovers several key insights, including a systematic evaluation of diverse techniques used in mobile GUI agents, optimal module configurations across model scales, the inherent limitations of current LFMs, and actionable guidelines for designing more capable and cost efficient mobile agents.

移动智能体评测基准模块化多路径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。