arXiv:2504.07981cs.CVcs.HC2025-04被引 268

针对专业高分辨率界面的智能识别难题,提出新基准与高效搜索方法。

ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

  • 基于专家标注的高分辨率界面图像构建跨领域评测基准
  • 现有模型在该数据集上最高准确率仅18.9%,新方法达48.1%
  • 无需额外训练,利用规划器知识引导分层搜索提升定位精度

多模态大语言模型(MLLM)在通用网页浏览和手机操作等任务中已取得显著进展,但在专业领域的应用仍不充分。这些专业工作流对图形用户界面感知模型带来独特挑战:高分辨率屏幕、目标区域小、环境复杂。本文提出ScreenSpot-Pro,一个用于严格评估MLLM在高分辨率专业场景下界面定位能力的新基准。该基准包含来自五个行业、三种操作系统共23个应用的真实高分辨率图像,并配有专家标注。现有界面定位模型在此数据集上表现不佳,最佳模型准确率仅为18.9%。实验表明,有策略地缩小搜索范围可提升准确率。基于此,我们提出ScreenSeekeR,一种利用强规划器的界面知识引导级联搜索的视觉搜索方法,在无需额外训练的情况下达到48.1%的准确率,实现当前最优性能。代码、数据与排行榜见https://gui-agent.github.io/grounding-leaderboard。

原文摘要 · Abstract (English)

Recent advancements in Multi-modal Large Language Models (MLLMs) have led to significant progress in developing GUI agents for general tasks such as web browsing and mobile phone use. However, their application in professional domains remains under-explored. These specialized workflows introduce unique challenges for GUI perception models, including high-resolution displays, smaller target sizes, and complex environments. In this paper, we introduce ScreenSpot-Pro, a new benchmark designed to rigorously evaluate the grounding capabilities of MLLMs in high-resolution professional settings. The benchmark comprises authentic high-resolution images from a variety of professional domains with expert annotations. It spans 23 applications across five industries and three operating systems. Existing GUI grounding models perform poorly on this dataset, with the best model achieving only 18.9%. Our experiments reveal that strategically reducing the search area enhances accuracy. Based on this insight, we propose ScreenSeekeR, a visual search method that utilizes the GUI knowledge of a strong planner to guide a cascaded search, achieving state-of-the-art performance with 48.1% without any additional training. We hope that our benchmark and findings will advance the development of GUI agents for professional applications. Code, data and leaderboard can be found at https://gui-agent.github.io/grounding-leaderboard.

GUI定位高分辨率大模型应用专业场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。