用局部区域+IoU优化,提升GUI元素定位精度
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
- 通过放大区域提案聚焦目标元素,减少干扰信息
- 使用IoU感知损失函数,使定位准确率提升13%
- 适合需要高精度界面操作的自动化任务研究者
面向图形用户界面(GUI)自动化的视觉智能体模型正快速发展,其核心挑战在于跨平台精准定位界面元素。现有仅依赖视觉的GUI智能体直接从大而杂乱的截图中进行定位,需处理大量无关信息,导致准确性下降。此外,这些方法通常采用基础交叉熵损失学习定位目标,难以有效捕捉定位质量,相比标准目标检测指标如交并比(IoU)表现较差。为此,我们提出R-VLM,一种新型GUI定位方法,利用缩放后的区域提案实现精确元素定位,并设计了基于IoU感知的目标函数,促进模型收敛到高IoU预测。该方法弥合了视觉语言模型与传统目标检测技术之间的差距,在ScreenSpot和AgentStudio两个基准上,跨多种GUI平台的定位准确率提升了13%。同时,在AITW和Mind2Web基准上的GUI导航任务中,准确率绝对提升达3.2%-9.7%。
原文摘要 · Abstract (English)
Visual agent models for automating human activities on Graphical User Interfaces (GUIs) have emerged as a promising research direction, driven by advances in large Vision Language Models (VLMs). A critical challenge in GUI automation is the precise grounding of interface elements across diverse platforms. Existing vision-only GUI agents directly ground elements from large and cluttered screenshots, requiring them to process substantial irrelevant information that compromises their accuracy. In addition, these approaches typically employ basic cross-entropy loss for learning grounding objectives, which fails to effectively capture grounding quality compared to established object detection metrics like Intersection-over-Union (IoU). To address these issues, we introduce R-VLM, a novel GUI grounding approach that leverages zoomed-in region proposals for precise element localization. We also propose an IoU-aware objective function that facilitates model convergence toward high IoU predictions. Our approach bridges the gap between VLMs and conventional object detection techniques, improving the state-of-the-art grounding accuracy by 13% across diverse GUI platforms on the GUI grounding benchmarks ScreenSpot and AgentStudio. In addition, our R-VLM approach shows 3.2-9.7% absolute accuracy improvements in GUI navigation tasks on the AITW and Mind2Web benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。