arXiv:2502.08226cs.CVcs.AI2025-02CVPR被引 8

无需训练的GUI理解框架,同时解决元素定位与描述问题。

TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents

  • 通过分层屏幕解析与空间增强描述模块,实现多粒度界面理解。
  • 在4个数据集上动作定位表现超越现有方法,界面描述任务优于ToL代理。
  • 适合需要跨平台通用性的GUI智能体开发者使用。

大型视觉语言模型(LVLM)的进步推动了基于LVLM的图形用户界面(GUI)智能体的发展。基于训练的方法如CogAgent和SeeClick因依赖特定数据集训练,难以实现跨数据集和跨平台泛化;通用型LVLM如GPT-4V采用标记集合(SoM)进行操作定位,但获取SoM标签需依赖HTML等元数据,而这些信息在不同平台上并不一致。此外,现有方法通常仅专注于单一GUI任务,而非实现全面的界面理解。为解决上述局限,我们提出TRISHUL——一种无需训练的智能体框架,可增强通用型LVLM以实现全面的GUI理解。与以往仅关注操作定位或界面指代的任务不同,TRISHUL无缝融合两者。其核心包括分层屏幕解析(HSP)与空间增强元素描述(SEED)模块,协同生成具有多粒度、空间与语义丰富性的界面元素表征。实验表明,TRISHUL在ScreenSpot、VisualWebBench、AITW和Mind2Web数据集上的动作定位表现更优;在界面指代任务中,其在ScreenPR基准上超越ToL代理,树立了鲁棒且适应性强的GUI理解新标准。

原文摘要 · Abstract (English)

Recent advancements in Large Vision Language Models (LVLMs) have enabled the development of LVLM-based Graphical User Interface (GUI) agents under various paradigms. Training-based approaches, such as CogAgent and SeeClick, struggle with cross-dataset and cross-platform generalization due to their reliance on dataset-specific training. Generalist LVLMs, such as GPT-4V, employ Set-of-Marks (SoM) for action grounding, but obtaining SoM labels requires metadata like HTML source, which is not consistently available across platforms. Moreover, existing methods often specialize in singular GUI tasks rather than achieving comprehensive GUI understanding. To address these limitations, we introduce TRISHUL, a novel, training-free agentic framework that enhances generalist LVLMs for holistic GUI comprehension. Unlike prior works that focus on either action grounding (mapping instructions to GUI elements) or GUI referring (describing GUI elements given a location), TRISHUL seamlessly integrates both. At its core, TRISHUL employs Hierarchical Screen Parsing (HSP) and the Spatially Enhanced Element Description (SEED) module, which work synergistically to provide multi-granular, spatially, and semantically enriched representations of GUI elements. Our results demonstrate TRISHUL's superior performance in action grounding across the ScreenSpot, VisualWebBench, AITW, and Mind2Web datasets. Additionally, for GUI referring, TRISHUL surpasses the ToL agent on the ScreenPR benchmark, setting a new standard for robust and adaptable GUI comprehension.

GUI理解视觉语言模型智能体界面定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。