从互联网视频自动构建大规模GUI操作数据集,提升AI助手泛化能力
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining

- 通过粗到精筛选策略,从无标注视频中提取高质量GUI操作轨迹
- 构建含1200万条轨迹的WildGUI数据集,覆盖1500+应用和网站
- 基于该数据集预训练模型,在多个任务上提升5%-20%,适合通用GUI代理研究
近年来,多模态大语言模型推动了图形用户界面(GUI)智能体的发展,但其泛化能力受限于大规模训练数据的缺乏。现有数据集依赖昂贵的人工标注,且局限于特定领域。为此,我们提出Video2GUI,一个完全自动化框架,可直接从无标注互联网视频中提取具有语义意义的GUI交互轨迹。该框架采用粗到精的过滤策略,识别高质量的GUI教程视频,并将其转换为结构化的智能体操作轨迹。通过对5亿条视频元数据进行处理,我们构建了WildGUI数据集,包含超过1200万条交互轨迹,覆盖1500多个应用与网站。在WildGUI上对Qwen2.5-VL和Mimo-VL进行预训练后,多项GUI定位与动作预测基准测试性能提升5%-20%,达到或超越当前最优水平。我们将公开WildGUI数据集与Video2GUI工具链,以支持未来GUI智能体研究。
原文摘要 · Abstract (English)
Recent advances in multimodal large language models have driven growing interest in graphical user interface (GUI) agents, yet their generalization remains constrained by the scarcity of large-scale training data spanning diverse real-world applications. Existing datasets rely heavily on costly manual annotations and are typically confined to narrow domains. To address this challenge, we propose Video2GUI, a fully automated framework that extracts grounded GUI interaction trajectories directly from unlabeled Internet videos. Video2GUI employs a coarse-to-fine filtering strategy to identify high-quality GUI tutorial videos and convert them into structured agent trajectories. Applying this pipeline to 500 million video metadata entries, we construct WildGUI, a large-scale dataset containing 12 million interaction trajectories spanning over 1,500 applications and websites. Pre-training Qwen2.5-VL and Mimo-VL on WildGUI yields consistent improvements of 5-20% across multiple GUI grounding and action benchmarks, matching or surpassing state-of-the-art performance. We will release both the WildGUI dataset and the Video2GUI pipeline to support future research of GUI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。