新基准MobileWorld挑战真实手机使用场景,提升任务复杂度与交互真实性。
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
- 构建跨应用长流程任务,平均步骤数达27.8,多应用任务占比62.2%
- 引入用户交互与MCP增强任务,真实模拟模糊指令与混合工具使用
- 基于开源替代方案实现可复现评估,支持精准验证与代码级控制
现有移动应用基准AndroidWorld虽具可复现性与确定性评估,但因当前智能体成功率超90%而趋于饱和。其缺乏电商、企业通信等关键应用类别,且未反映真实使用中模糊指令与多工具混合的场景。为此,我们提出MobileWorld,涵盖20个应用的201项任务,强调长周期跨应用工作流,平均完成步骤27.8(远高于AndroidWorld的14.3),多应用任务占比62.2%(原为9.5%)。通过采用开源替代方案(如用Mattermost替代Slack),实现可观察、可控制的环境,支持源码修改与数据库直接访问以精确验证。新增代理-用户交互与模型上下文协议(MCP)增强任务,评估代理在用户感知与混合工具场景下的表现。我们开发了支持用户交互与MCP调用的规划-执行框架。实验显示,最优代理框架成功率仅51.7%,端到端模型为20.9%,显著低于AndroidWorld,凸显未来研究的巨大空间。
原文摘要 · Abstract (English)
Among existing online mobile-use benchmarks, AndroidWorld has emerged as the dominant benchmark due to its reproducible environment and deterministic evaluation; however, recent agents achieving over 90% success rates indicate its saturation and motivate the need for a more challenging benchmark. In addition, its environment lacks key application categories, such as e-commerce and enterprise communication, and does not reflect realistic mobile-use scenarios characterized by vague user instructions and hybrid tool usage. We introduce MobileWorld, a substantially more challenging benchmark designed to reflect real-world usage through 201 tasks across 20 applications. MobileWorld derives its difficulty from an emphasis on long-horizon, cross-application workflows, requiring nearly twice as many completion steps on average (27.8 vs. 14.3) and featuring a significantly higher proportion of multi-app tasks (62.2% vs. 9.5%) than AndroidWorld. To overcome the limitations of existing environments, MobileWorld achieves a balance between production-grade utility and reproducible evaluation by utilizing open-source alternatives to industry standards (e.g., Mattermost for Slack). This approach enables a fully observable and controlled environment through source code modification and direct backend database access for precise verification. MobileWorld also introduces novel task categories, including agent-user interaction and Model Context Protocol (MCP)-augmented tasks, for evaluating agents in user-aware, hybrid-tool scenarios. To facilitate evaluation, we develop a planner-executor agentic framework with extended action spaces to support user interactions and MCP calls. Our results reveal a sharp performance drop compared to AndroidWorld, with the best agentic framework and end-to-end model achieving 51.7% and 20.9% success rates, respectively, highlighting ample headroom for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。