首个生成应用界面图像的视觉世界模型,提升手机自动化代理的长程规划能力。
ViMo: A Generative Visual GUI World Model for App Agents
- 将界面生成拆分为图形与文字两部分,用符号占位符保留视觉结构
- 在真实应用上生成可读性强、功能准确的未来界面图像,支持决策预测
- 适合需要复杂交互规划的移动应用自动化研究者使用
应用代理通过图形用户界面(GUI)自主操作移动端应用,已在实际应用中引发广泛关注。然而,它们在长程任务规划中常表现不佳,难以找到复杂任务中的最优动作序列。为此,世界模型可通过用户动作预测未来的GUI观察,从而提升代理规划效率。但现有世界模型主要生成文本描述,缺乏关键的视觉细节。为填补这一空白,我们提出ViMo,首个专为生成未来应用界面图像而设计的视觉世界模型。针对图像块中文字生成易受像素误差影响的问题,我们将界面生成分解为图形与文字内容生成。提出符号化文本表示(STR),在保留图形的同时以符号占位符叠加文本内容。基于此设计,ViMo采用STR预测器生成未来界面的图形,并通过GUI-Text预测器生成对应文字。此外,我们利用ViMo增强代理任务,预测不同动作选项的结果。实验表明,ViMo能生成视觉合理且功能有效的界面,使应用代理做出更明智决策。
原文摘要 · Abstract (English)
App agents, which autonomously operate mobile Apps through Graphical User Interfaces (GUIs), have gained significant interest in real-world applications. Yet, they often struggle with long-horizon planning, failing to find the optimal actions for complex tasks with longer steps. To address this, world models are used to predict the next GUI observation based on user actions, enabling more effective agent planning. However, existing world models primarily focus on generating only textual descriptions, lacking essential visual details. To fill this gap, we propose ViMo, the first visual world model designed to generate future App observations as images. For the challenge of generating text in image patches, where even minor pixel errors can distort readability, we decompose GUI generation into graphic and text content generation. We propose a novel data representation, the Symbolic Text Representation~(STR) to overlay text content with symbolic placeholders while preserving graphics. With this design, ViMo employs a STR Predictor to predict future GUIs' graphics and a GUI-text Predictor for generating the corresponding text. Moreover, we deploy ViMo to enhance agent-focused tasks by predicting the outcome of different action options. Experiments show ViMo's ability to generate visually plausible and functionally effective GUIs that enable App agents to make more informed decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。