arXiv:2508.09241cs.CV2025-08被引 6

首个评估GUI代理细粒度状态控制能力的基准,揭示视觉定位是当前模型瓶颈。

FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents

  • 构建多平台细粒度任务基准,含2257个任务分四类评估
  • 顶尖模型细粒度交互准确率仅32.8%,视觉定位提升14.9%成功率
  • 提出可插拔诊断工具VDA,首次量化分析感知与定位能力

随着生成式人工智能的发展,图形用户界面(GUI)代理在自然语言指令下自主完成日常任务展现出巨大潜力。然而,现有评估框架存在根本缺陷:过度关注粗粒度任务完成,忽视真实应用中至关重要的细粒度控制能力。为此,我们提出FineState-Bench,首个针对细粒度GUI代理操作的评估与诊断标准,用于量化细粒度控制能力。该多平台(桌面、Web、移动)框架包含2257个任务,分为四个组件,并采用四阶段指标实现从感知到控制的全面评估。为分析精细化操作中的感知与定位能力,我们开发了即插即用的视觉诊断助手(VDA),实现首次定量解耦分析。在本基准上的实验表明,最先进模型的细粒度交互准确率仅为32.8%。通过受控实验使用VDA量化视觉能力影响,发现理想视觉定位可使Gemini-2.5-Flash的成功率提升14.9%。我们的诊断框架首次确认,当前GUI代理的主要瓶颈在于基础视觉定位能力。所有资源均开源。

原文摘要 · Abstract (English)

With the rapid advancement of generative artificial intelligence technology, Graphical User Interface (GUI) agents have demonstrated tremendous potential for autonomously managing daily tasks through natural language instructions. However, current evaluation frameworks for GUI agents suffer from fundamental flaws: existing benchmarks overly focus on coarse-grained task completion while neglecting fine-grained control capabilities crucial for real-world applications. To address this, we introduce FineState-Bench, the first evaluation and diagnostic standard for fine-grained GUI proxy operations, designed to quantify fine-grained control. This multi-platform (desktop, Web, mobile) framework includes 2257 task benchmarks in four components and uses a four-phase indicator for comprehensive perception-to-control assessment. To analyze perception and positioning for refined operations, we developed the plug-and-play Visual Diagnostic Assistant (VDA), enabling the first quantitative decoupling analysis of these capabilities. Experimental results on our benchmark show that the most advanced models achieve only 32.8% fine-grained interaction accuracy. Using our VDA in controlled experiments, quantifying the impact of visual capabilities, we showed that ideal visual localization boosts Gemini-2.5-Flash's success rate by 14.9\%. Our diagnostic framework confirms for the first time that the primary bottleneck for current GUI proxies is basic visual positioning capability.All resources are fully open-source. github: https://github.com/AnonymousThewarehouse/FineState-Bench huggingface: https://huggingface.co/datasets/Willtime2006/Static-FineBench

GUI代理细粒度控制评估基准视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。