用网页自动生成多粒度数据,提升视觉语言模型的界面理解能力
EDGE: Enhanced Grounded GUI Understanding with Enriched Multi-Granularity Synthetic Data
- 自动从网页生成大规模多粒度合成数据
- 在多个基准上显著提升界面理解性能
- 适合需要免标注训练数据的GUI智能体研究者
在各类应用的图形用户界面(GUI)上运行的自主代理具有巨大实用价值。与依赖结构化文本和定制后端的大语言模型方法不同,使用大视觉语言模型(LVLMs)的方法更直观且适应性强,因其可直接视觉感知并操作屏幕,在无文本元数据和专用后端的通用场景中不可或缺。由于现有研究缺乏高质量的GUI任务训练数据,本文提出一种数据驱动方法,通过EDGE框架自动从网络网页生成大规模、多粒度的训练数据,以增强LVLM在GUI理解与交互方面的能力。在多个GUI与智能体基准上的评估结果表明,使用EDGE生成的数据训练的模型展现出更强的网页理解能力,并可轻松迁移到此前未见过的桌面和移动端环境。该方法显著降低了对人工标注的依赖,使研究人员能利用网络上丰富的公开资源推进工作。源代码、数据集及模型已公开于https://anonymous.4open.science/r/EDGE-1CDB。
原文摘要 · Abstract (English)
Autonomous agents operating on the graphical user interfaces (GUIs) of various applications hold immense practical value. Unlike the large language model (LLM)-based methods which rely on structured texts and customized backends, the approaches using large vision-language models (LVLMs) are more intuitive and adaptable as they can visually perceive and directly interact with screens, making them indispensable in general scenarios without text metadata and tailored backends. Given the lack of high-quality training data for GUI-related tasks in existing work, this paper aims to enhance the GUI understanding and interacting capabilities of LVLMs through a data-driven approach. We propose EDGE, a general data synthesis framework that automatically generates large-scale, multi-granularity training data from webpages across the Web. Evaluation results on various GUI and agent benchmarks demonstrate that the model trained with the dataset generated through EDGE exhibits superior webpage understanding capabilities, which can then be easily transferred to previously unseen desktop and mobile environments. Our approach significantly reduces the dependence on manual annotations, empowering researchers to harness the vast public resources available on the Web to advance their work. Our source code, the dataset and the model are available at https://anonymous.4open.science/r/EDGE-1CDB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。