arXiv:2410.23218cs.CLcs.CV2024-10被引 360

开源的通用图形界面智能体模型,提升跨平台操作能力。

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

  • 构建跨平台GUI数据合成工具,生成超1300万元素数据集
  • 在六个基准上超越现有最优模型,尤其在未知界面表现优异
  • 适合研究开源视觉语言模型与自动化交互的开发者

当前构建图形界面智能体多依赖GPT-4o、GeminiProVision等商用视觉语言模型。由于开源模型在界面定位和分布外(OOD)场景下性能差距显著,研究者普遍不愿使用。为此,我们开发了OS-Atlas——一个基础性图形界面动作模型,通过数据与建模双重创新,在界面定位和分布外智能体任务中表现突出。我们投入大量工程资源,构建了一个开源工具链,可跨Windows、Linux、macOS、Android及网页平台合成界面定位数据。基于此,我们发布了迄今最大的开源跨平台界面定位语料库,包含超过1300万图形界面元素。该数据集结合模型训练优化,使OS-Atlas能准确理解界面截图并泛化至未见界面。在覆盖移动、桌面、网页三类平台的六个基准上评估,其性能显著优于先前最优模型。评估还揭示了持续提升和扩展开源视觉语言模型智能体能力的关键洞见。

原文摘要 · Abstract (English)

Existing efforts in building GUI agents heavily rely on the availability of robust commercial Vision-Language Models (VLMs) such as GPT-4o and GeminiProVision. Practitioners are often reluctant to use open-source VLMs due to their significant performance lag compared to their closed-source counterparts, particularly in GUI grounding and Out-Of-Distribution (OOD) scenarios. To facilitate future research in this area, we developed OS-Atlas - a foundational GUI action model that excels at GUI grounding and OOD agentic tasks through innovations in both data and modeling. We have invested significant engineering effort in developing an open-source toolkit for synthesizing GUI grounding data across multiple platforms, including Windows, Linux, MacOS, Android, and the web. Leveraging this toolkit, we are releasing the largest open-source cross-platform GUI grounding corpus to date, which contains over 13 million GUI elements. This dataset, combined with innovations in model training, provides a solid foundation for OS-Atlas to understand GUI screenshots and generalize to unseen interfaces. Through extensive evaluation across six benchmarks spanning three different platforms (mobile, desktop, and web), OS-Atlas demonstrates significant performance improvements over previous state-of-the-art models. Our evaluation also uncovers valuable insights into continuously improving and scaling the agentic capabilities of open-source VLMs.

GUI智能体开源模型跨平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。