arXiv:2512.14014cs.AI2025-12被引 13

用自然语言描述界面状态,提升移动端智能体的规划能力。

MobileWorldBench: Towards Semantic World Modeling For Mobile Agents

  • 以自然语言替代像素预测,构建语义世界模型
  • 在140万样本数据集上提升智能体任务成功率
  • 适合做移动端自动化与AI代理研究的开发者

世界模型在提升具身智能体任务表现方面展现出巨大潜力。以往工作多聚焦于像素空间的世界模型,但在GUI场景中,未来状态的复杂视觉元素预测常面临实际困难。本文探索了一种面向移动端智能体的替代性世界建模方式:用自然语言描述状态转移,而非预测原始像素。首先,提出MobileWorldBench基准,评估视觉语言模型(VLM)作为移动GUI智能体世界模型的能力;其次,发布MobileWorld大规模数据集,包含140万样本,显著提升VLM的世界建模能力;最后,设计一种新框架,将VLM世界模型融入移动智能体的规划流程,实证表明语义世界模型能直接提升任务成功率达显著水平。代码与数据集已开源。

原文摘要 · Abstract (English)

World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face practical limitations in GUI settings, where predicting complex visual elements in future states is often difficult. In this work, we explore an alternative formulation of world modeling for GUI agents, where state transitions are described in natural language rather than predicting raw pixels. First, we introduce MobileWorldBench, a benchmark that evaluates the ability of vision-language models (VLMs) to function as world models for mobile GUI agents. Second, we release MobileWorld, a large-scale dataset consisting of 1.4M samples, that significantly improves the world modeling capabilities of VLMs. Finally, we propose a novel framework that integrates VLM world models into the planning framework of mobile agents, demonstrating that semantic world models can directly benefit mobile agents by improving task success rates. The code and dataset is available at https://github.com/jacklishufan/MobileWorld

世界模型移动端代理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。