arXiv:2509.04243cs.CVcs.AI2025-09被引 3

让视觉语言模型学会主动看界面,精准定位元素。

Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding

  • 用自进化偏好优化方法,逐步提升模型多步感知能力。
  • 在ScreenSpot-Pro上达55.7分,7B模型新纪录。
  • 适合需要高精度界面理解的AI交互研究者。

视觉语言模型(VLMs)在连接视觉感知与语言推理方面取得了显著进展。OpenAI o3模型引入的缩放搜索策略有效激发了VLMs的主动感知能力,提升了下游任务表现。然而,在高分辨率输入和复杂多元素视觉交互下,使VLMs对合适图像区域进行有效推理仍是GUI定位的核心挑战。本文提出LASER框架,通过自进化方式逐步赋予VLMs多步感知能力,实现精确坐标预测。该方法结合蒙特卡洛质量评估与基于交并比(IoU)的区域质量评估,共同促进高质量偏好数据的构建,明确引导模型关注指令相关关键区域,并根据任务复杂度自适应分配推理步骤。在ScreenSpot Pro和ScreenSpot-v2基准上的全面实验验证了方法的有效性。此外,在GTA1-7B上微调后,LASER在ScreenSpot-Pro上取得55.7分,成为7B规模模型中的最新最优(SoTA)。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) have recently achieved significant progress in bridging visual perception and linguistic reasoning. Recently, OpenAI o3 model introduced a zoom-in search strategy that effectively elicits active perception capabilities in VLMs, improving downstream task performance. However, enabling VLMs to reason effectively over appropriate image regions remains a core challenge in GUI grounding, particularly under high-resolution inputs and complex multi-element visual interactions. In this work, we propose LASER, a self-evolving framework that progressively endows VLMs with multi-step perception capabilities, enabling precise coordinate prediction. Specifically, our approach integrate Monte Carlo quality estimation with Intersection-over-Union (IoU)-based region quality evaluation to jointly encourage both accuracy and diversity in constructing high-quality preference data. This combination explicitly guides the model to focus on instruction-relevant key regions while adaptively allocating reasoning steps based on task complexity. Comprehensive experiments on the ScreenSpot Pro and ScreenSpot-v2 benchmarks demonstrate consistent performance gains, validating the effectiveness of our method. Furthermore, when fine-tuned on GTA1-7B, LASER achieves a score of 55.7 on the ScreenSpot-Pro benchmark, establishing a new state-of-the-art (SoTA) among 7B-scale models.

GUI定位视觉语言模型主动感知偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。