自动生成桌面界面描述数据,推动智能助手理解视觉元素
DeskVision: Large Scale Desktop Region Captioning for Advanced GUI Agents
- 用自动化工具生成带丰富描述的桌面界面数据
- 构建了超大规模桌面界面数据集,支持多种系统和控件
- 新模型在无需复杂结构下实现顶尖视觉理解性能
图形用户界面(GUI)数据匮乏严重制约了桌面场景下智能助手的发展。为此,我们提出自动化数据生成流水线AutoCaptioner,可低人力成本生成丰富描述数据。基于此,我们构建了新型大规模桌面界面数据集DeskVision及最大规模测试基准DeskVision-Eval,覆盖日常使用场景、多样系统与界面元素,并附带详细描述。基于DeskVision,我们训练出新模型GUIExplorer,其在视觉元素理解与定位任务中达到当前最优表现,且无需复杂架构设计。我们通过消融实验验证了DeskVision对多种大视觉语言模型的有效性。我们认为AutoCaptioner与DeskVision将显著推动GUI助手发展,相关资源将开源共享。
原文摘要 · Abstract (English)
The limitation of graphical user interface (GUI) data has been a significant barrier to the development of GUI agents today, especially for the desktop / computer use scenarios. To address this, we propose an automated GUI data generation pipeline, AutoCaptioner, which generates data with rich descriptions while minimizing human effort. Using AutoCaptioner, we created a novel large-scale desktop GUI dataset, DeskVision, along with the largest desktop test benchmark, DeskVision-Eval, which reflects daily usage and covers diverse systems and UI elements, each with rich descriptions. With DeskVision, we train a new GUI understanding model, GUIExplorer. Results show that GUIExplorer achieves state-of-the-art (SOTA) performance in understanding/grounding visual elements without the need for complex architectural designs. We further validated the effectiveness of the DeskVision dataset through ablation studies on various large visual language models (LVLMs). We believe that AutoCaptioner and DeskVision will significantly advance the development of GUI agents, and will open-source them for the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。