arXiv:2410.18967cs.CVcs.CL2024-10ICLR被引 58

Ferret-UI 2能跨平台理解界面,支持多设备精准操作。

Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms

  • 通过自适应缩放实现高分辨率界面感知
  • 用GPT-4o生成9个子任务×5平台的训练数据
  • 在跨平台任务中表现优异,适合多端交互开发

构建通用用户界面(UI)理解模型面临平台多样性、分辨率差异和数据有限等基础挑战。本文提出Ferret-UI 2,一种面向跨平台通用界面理解的多模态大语言模型,覆盖iPhone、Android、iPad、Webpage及AppleTV等多种平台。在Ferret-UI基础上,Ferret-UI 2引入三项关键创新:支持多平台类型、通过自适应缩放实现高分辨率感知,以及基于GPT-4o与集合标记视觉提示的先进任务训练数据生成。这些改进使Ferret-UI 2能够执行复杂、以用户为中心的交互任务,具备高度灵活性与适应性。在包括9个子任务×5平台的指代与定位任务、GUIDE下一步动作预测数据集以及GUI-World跨平台基准上的大量实验证明,Ferret-UI 2显著优于Ferret-UI,且展现出强大的跨平台迁移能力。

原文摘要 · Abstract (English)

Building a generalist model for user interface (UI) understanding is challenging due to various foundational issues, such as platform diversity, resolution variation, and data limitation. In this paper, we introduce Ferret-UI 2, a multimodal large language model (MLLM) designed for universal UI understanding across a wide range of platforms, including iPhone, Android, iPad, Webpage, and AppleTV. Building on the foundation of Ferret-UI, Ferret-UI 2 introduces three key innovations: support for multiple platform types, high-resolution perception through adaptive scaling, and advanced task training data generation powered by GPT-4o with set-of-mark visual prompting. These advancements enable Ferret-UI 2 to perform complex, user-centered interactions, making it highly versatile and adaptable for the expanding diversity of platform ecosystems. Extensive empirical experiments on referring, grounding, user-centric advanced tasks (comprising 9 subtasks $\times$ 5 platforms), GUIDE next-action prediction dataset, and GUI-World multi-platform benchmark demonstrate that Ferret-UI 2 significantly outperforms Ferret-UI, and also shows strong cross-platform transfer capabilities.

界面理解多模态模型跨平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。