构建复杂真实场景的移动端智能体评测基准,揭示现有模型在高阶任务中表现不足。
Benchmarking General Mobile Assistants in Challenging Real-World Scenarios

- 设计7个开源应用与300项多层级任务,覆盖从简单操作到复杂流程。
- 8个前沿模型在复杂任务上性能显著下降,无法可靠满足真实用户需求。
- 验证了上下文保持与状态追踪等框架设计对复杂任务的关键提升作用。
图形用户界面已成为评估多模态交互任务中自主智能体的重要环境。现有基准如AndroidWorld和MobileWorld为移动端智能体评估奠定了坚实基础,但其应用覆盖范围和任务设计尚未充分反映真实移动使用中的多样性和复杂性。我们提出GMA,一个面向通用移动端助手在挑战性真实场景中的评测基准。GMA基于开源项目构建7个应用,涵盖生活方式分享、旅行规划等领域的300项任务,分为四个难度层级,从原子操作到复杂多步工作流。我们评估了8个前沿模型,发现随着任务复杂度增加,性能显著下降,当前智能体仍远未达到可靠处理真实用户需求的水平。进一步在统一环境、模型设置和任务分类下,开展受控消融实验,考察上下文保留与显式状态追踪等代理框架选择的影响。结果表明,合适的框架设计可显著提升性能,尤其在高难度工作流中,但不同基础模型对具体设计的有效性存在差异。总体而言,GMA通过扩展应用覆盖范围和任务复杂度,补充了现有基准,为移动智能体评估及框架设计如何支持复杂工作流可靠执行提供了有力测试平台。
原文摘要 · Abstract (English)
Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks. Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design do not yet fully capture the diversity and complexity of realistic mobile use. We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios. GMA introduces seven applications based on open-source projects, spanning domains such as lifestyle sharing and travel planning, and 300 tasks across four difficulty tiers, from atomic actions to complex multi-step workflows. We evaluate eight frontier models and find that performance declines substantially as task complexity increases, with current agents remaining far from reliably handling realistic user requirements. We further conduct controlled ablation studies of agent harness choices, including context retention and explicit state tracking, under a shared environment, model setting, and task taxonomy. Results show that appropriate harness design can meaningfully improve performance, particularly on demanding workflows, while the effectiveness of specific designs can vary across foundation models. Overall, GMA complements existing benchmarks by expanding application coverage and task complexity, providing a challenging testbed for evaluating mobile agents and studying how harness design supports reliable execution in complex mobile workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。