用网页教程训练通用图形界面智能体,数据量超14万条。
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
- 从多模态网络教程自动构建界面操作轨迹数据
- 生成14.3万条跨系统、200+应用的轨迹数据集
- 基于大模型微调,性能比基线提升约10%
构建图形用户界面(GUI)智能体是模拟人类与计算机或手机交互以完成多样化任务的有前景方向。然而,开发通用GUI智能体的主要挑战在于缺乏跨操作系统和应用的充足轨迹数据,主要由于人工标注成本高昂。本文提出TongUI框架,通过学习丰富的多模态网络教程来构建通用GUI智能体。具体而言,我们爬取并处理在线GUI教程(如视频和文章),将其转化为GUI智能体轨迹数据,从而生成包含14.3万条轨迹数据的GUI-Net数据集,覆盖五个操作系统和200多个应用。我们基于GUI-Net对Qwen2.5-VL-3B/7B模型进行微调,得到TongUI智能体,在常用定位与导航基准测试中表现显著提升,多个基准上性能优于基线约10%,验证了GUI-Net数据集的有效性,并凸显了TongUI框架的重要性。代码、GUI-Net数据集及训练模型将很快全部开源。
原文摘要 · Abstract (English)
Building Graphical User Interface (GUI) agents is a promising research direction, which simulates human interaction with computers or mobile phones to perform diverse GUI tasks. However, a major challenge in developing generalized GUI agents is the lack of sufficient trajectory data across various operating systems and applications, mainly due to the high cost of manual annotations. In this paper, we propose the TongUI framework that builds generalized GUI agents by learning from rich multimodal web tutorials. Concretely, we crawl and process online GUI tutorials (such as videos and articles) into GUI agent trajectory data, through which we produce the GUI-Net dataset containing 143K trajectory data across five operating systems and more than 200 applications. We develop the TongUI agent by fine-tuning Qwen2.5-VL-3B/7B models on GUI-Net, which show remarkable performance improvements on commonly used grounding and navigation benchmarks, outperforming baseline agents about 10\% on multiple benchmarks, showing the effectiveness of the GUI-Net dataset and underscoring the significance of our TongUI framework. We will fully open-source the code, the GUI-Net dataset, and the trained models soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。