用屏幕图像模拟手机界面动态变化,提升AI交互训练效率。
UISim: An Interactive Image-Based UI Simulator for Dynamic Mobile Environments
- 基于屏幕图像预测下一界面布局并合成真实视觉图像
- 生成的界面过渡更真实连贯,优于现有端到端方法
- 适合做UI测试、原型快速开发和AI导航任务训练
由于真实移动环境的动态性和多样性,开发和测试用户界面(UI)以及训练智能体与之交互面临挑战。现有方法多依赖繁琐的物理设备或对截图的静态分析,难以实现可扩展测试和智能体训练。我们提出UISim,一种基于图像的新型交互式UI模拟器,仅通过屏幕图像即可在动态移动环境中进行探索。系统采用两阶段方法:给定初始屏幕图像和用户操作后,先预测下一状态的抽象布局,再据此合成视觉一致的新图像。该方法实现了真实的界面状态转换模拟。UISim在UI测试、快速原型设计和合成数据生成方面已带来即时效益,并为智能体的界面导航任务规划等高级应用铺平道路。实验表明,与端到端界面生成基线相比,UISim在生成真实且连贯的后续界面状态方面表现更优,体现了其高保真度和在简化界面开发、增强智能体训练方面的潜力。
原文摘要 · Abstract (English)
Developing and testing user interfaces (UIs) and training AI agents to interact with them are challenging due to the dynamic and diverse nature of real-world mobile environments. Existing methods often rely on cumbersome physical devices or limited static analysis of screenshots, which hinders scalable testing and the development of intelligent UI agents. We introduce UISim, a novel image-based UI simulator that offers a dynamic and interactive platform for exploring mobile phone environments purely from screen images. Our system employs a two-stage method: given an initial phone screen image and a user action, it first predicts the abstract layout of the next UI state, then synthesizes a new, visually consistent image based on this predicted layout. This approach enables the realistic simulation of UI transitions. UISim provides immediate practical benefits for UI testing, rapid prototyping, and synthetic data generation. Furthermore, its interactive capabilities pave the way for advanced applications, such as UI navigation task planning for AI agents. Our experimental results show that UISim outperforms end-to-end UI generation baselines in generating realistic and coherent subsequent UI states, highlighting its fidelity and potential to streamline UI development and enhance AI agent training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。