arXiv:2411.13591cs.CVcs.AI2024-11被引 12

通过迭代缩小机制提升视觉语言模型的界面理解能力

Improved GUI Grounding via Iterative Narrowing

  • 设计视觉提示框架,用迭代方式逐步缩小界面元素范围
  • 在多平台界面数据集上显著优于基线模型
  • 适用于通用与微调后的视觉语言模型,提升零样本界面定位

图形用户界面(GUI)定位在增强视觉-语言模型(VLM)智能体能力方面起着关键作用。尽管通用VLM如GPT-4V在各类任务中表现强劲,但在GUI定位任务上的性能仍不理想。近期研究聚焦于针对零样本GUI定位对这些模型进行微调,取得了显著改进。本文提出一种视觉提示框架,采用迭代缩小机制,进一步提升通用及微调后模型在GUI定位中的表现。我们在涵盖多种用户界面平台的综合基准上评估该方法,并公开代码以支持结果复现。

原文摘要 · Abstract (English)

Graphical User Interface (GUI) grounding plays a crucial role in enhancing the capabilities of Vision-Language Model (VLM) agents. While general VLMs, such as GPT-4V, demonstrate strong performance across various tasks, their proficiency in GUI grounding remains suboptimal. Recent studies have focused on fine-tuning these models specifically for zero-shot GUI grounding, yielding significant improvements over baseline performance. We introduce a visual prompting framework that employs an iterative narrowing mechanism to further improve the performance of both general and fine-tuned models in GUI grounding. For evaluation, we tested our method on a comprehensive benchmark comprising various UI platforms and provided the code to reproduce our results.

GUI定位视觉语言模型迭代优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。