构建31万帧跨平台手机操作数据集,提升AI导航泛化能力
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
- 用YouTube视频自动构建带标注的跨平台手机操作数据集
- 模型在未见系统上性能提升18.11个百分点
- 支持持续更新,适合做移动智能体研究的团队使用
大型语言模型和视觉语言模型的发展推动了图形用户界面智能体的研究。我们提出MONDAY(从YouTube获取的移动操作系统导航任务数据集),包含20,000个教学视频中提取的313,000帧标注图像,覆盖多个平台的真实手机操作场景。在预训练中加入MONDAY的数据,模型展现出更强的跨平台泛化能力,在未见过的移动操作系统上平均性能提升18.11个百分点,优于仅在单一系统数据上训练的模型。为实现数据集随平台演进持续扩展,我们提出自动化框架:基于高精度OCR的场景检测(F1得分95.04%)、近乎完美的UI元素识别(命中率99.87%),以及创新的多步动作识别技术,可在不同界面布局中可靠提取操作序列。我们公开MONDAY数据集及自动化采集框架,以推动移动端导航智能体研究。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have sparked significant interest in developing GUI visual agents. We introduce MONDAY (Mobile OS Navigation Task Dataset for Agents from YouTube), a large-scale dataset of 313K annotated frames from 20K instructional videos capturing diverse real-world mobile OS navigation across multiple platforms. Models that include MONDAY in their pre-training phases demonstrate robust cross-platform generalization capabilities, consistently outperforming models trained on existing single OS datasets while achieving an average performance gain of 18.11%p on an unseen mobile OS platform. To enable continuous dataset expansion as mobile platforms evolve, we present an automated framework that leverages publicly available video content to create comprehensive task datasets without manual annotation. Our framework comprises robust OCR-based scene detection (95.04% F1score), near-perfect UI element detection (99.87% hit ratio), and novel multi-step action identification to extract reliable action sequences across diverse interface configurations. We contribute both the MONDAY dataset and our automated collection framework to facilitate future research in mobile OS navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。