用视觉语言模型理解手机操作行为,实现自动化的任务复现。
GUITrans2Act: Understanding User Operational Behaviors from Mobile GUI Interactions with Vision-Language Models

- 从操作视频中提取关键帧,生成分步操作指令。
- 在中文移动端基准测试中表现优于现有模型,任务成功率显著提升。
- 适合自动化测试、智能助手等需要理解用户操作的场景。
理解移动设备上的数字世界正从静态界面感知转向动态操作理解。该能力使模型能将视觉状态变化转化为操作知识,即描述操作类型、目标控件、文本参数和执行顺序的自然语言短句。然而,由于应用间界面设计高度多样且异构,现有视觉语言模型(VLM)难以准确推断这些底层操作。为此,我们提出Teach VLM,一种核心模型,通过从演示视频中提取并分析操作相关关键帧,将屏幕轨迹转化为分步操作知识。为解决标注数据稀缺问题,我们构建了系统化数据飞轮以实现可扩展的数据获取。我们进一步引入首个细粒度中文移动端屏幕教学基准(Chinese Mobile Screen Teach Benchmark)。基于Teach VLM,我们提出Teach-and-Repeat范式,利用生成的操作知识作为可解释的程序参考,指导下游屏幕执行代理。大量实验表明,Teach VLM显著优于强基线模型,在操作语义预测上达到当前最佳性能。此外,在Android World中的实验显示,该范式持续提升下游代理的任务成功率。Teach VLM与Teach-and-Repeat范式共同提供了一条从原始演示到可复用任务自动化的实用路径。
原文摘要 · Abstract (English)
Understanding the digital world on mobile devices is shifting from static UI perception to dynamic action comprehension. This capability enables models to convert visual state transitions into operational knowledge, defined as short natural-language sentences that describe action types, target UI elements, textual arguments, and execution orders. However, due to the highly diverse and heterogeneous UI designs across applications, existing vision-language models (VLMs) struggle to accurately infer these underlying operations. To bridge this gap, we introduce Teach VLM, a core model designed to translate mobile screen trajectories into step-wise operational knowledge by extracting and analyzing operation-related keyframes from demonstration videos. To address the scarcity of aligned training data, we develop a systematic data flywheel for scalable data acquisition. We further introduce a novel Chinese Mobile Screen Teach Benchmark for fine-grained evaluation. Building upon Teach VLM, we propose the Teach-and-Repeat paradigm, where the generated operational knowledge serves as an interpretable procedural reference to guide downstream screen-based execution agents. Extensive evaluations demonstrate that Teach VLM significantly outperforms strong VLM baselines, achieving state-of-the-art performance in operation semantics prediction. Furthermore, experiments in Android World show that our paradigm yields consistent Task Success Rate improvements for downstream agents. Together, Teach VLM and the Teach-and-Repeat paradigm offer a practical pathway from raw demonstrations to reusable task automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。