arXiv:2604.27974cs.CVcs.DB2026-04ACL被引 5

构建细粒度界面状态交互评估基准,精准定位失败原因。

FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting

论文配图:FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting
图 1 · 摘自论文原文
  • 设计多平台细粒度状态指令基准,明确目标状态
  • 提出四阶段评估指标,最高准确率仅32.8%
  • 引入可视化诊断助手,提升14.9%状态命中率

尽管大视觉语言模型进展迅速,细粒度状态条件化界面交互仍具挑战。现有评估覆盖有限、目标状态定义模糊,且过度依赖最终任务成功,难以揭示失败环节。为此,我们提出 FineState-Bench,一个评估智能体是否能正确将指令映射至目标控件并达到精确目标状态的基准。该基准包含跨桌面、网页和移动端的2,209个实例,涵盖四种交互类型和23种UI组件,每个实例均明确定义精确目标状态。我们进一步提出 FineState-Metrics 四阶段诊断流程,包括定位成功率(SR@Loc)、交互成功率(SR@Int)、定位后精确状态成功率(ES-SR@Loc)和交互后精确状态成功率(ES-SR@Int),并设计即插即用的 Visual Diagnostic Assistant(VDA),通过有无提示对比生成描述与边界框提示,诊断视觉定位原因。在 FineState-Bench 上,精确目标状态成功率仍低:网页平台最高为32.8%,跨平台平均仅22.8%。使用 VDA 提示后,Gemini-2.5-Flash 的 ES-SR@Int 提升14.9个百分点,表明视觉定位仍有巨大改进空间,但整体准确率仍不足以支撑可靠的细粒度状态条件交互。

原文摘要 · Abstract (English)

Despite the rapid progress of large vision-language models (LVLMs), fine-grained, state-conditioned GUI interaction remains challenging. Current evaluations offer limited coverage, imprecise target-state definitions, and an overreliance on final-task success, obscuring where and why agents fail. To address this gap, we introduce \textbf{FineState-Bench}, a benchmark that evaluates whether an agent can correctly ground an instruction to the intended UI control and reach the exact target state. FineState-Bench comprises 2,209 instances across desktop, web, and mobile platforms, spanning four interaction families and 23 UI component types, with each instance explicitly specifying an exact target state for fine-grained state setting. We further propose \textit{FineState-Metrics}, a four-stage diagnostic pipeline with stage-wise success rates: Localization Success Rate (SR@Loc), Interaction Success Rate (SR@Int), Exact State Success Rate at Locate (ES-SR@Loc), and Exact State Success Rate at Interact (ES-SR@Int), and a plug-and-play \textit{Visual Diagnostic Assistant} (VDA) that generates a Description and a bounding-box Localization Hint to diagnose visual grounding reason via controlled w/ vs.\ w/o comparisons. On FineState-Bench, exact goal-state success remains low: ES-SR@Int peaks at 32.8\% on Web and 22.8\% on average across platforms. With VDA localization hints, Gemini-2.5-Flash gains +14.9 ES-SR@Int points, suggesting substantial headroom from improved visual grounding, yet overall accuracy is still insufficient for reliable fine-grained state-conditioned interaction \href{https://github.com/FengxianJi/FineState-Bench}{Github.}

GUI交互评估基准视觉定位状态设置

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。