arXiv:2604.14262cs.LGcs.AI2026-04被引 1

GUI模型在空间推理任务中准确率暴跌,新框架揭示其系统性脆弱

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models

  • 通过独立扰动界面与指令,量化模型鲁棒性
  • 空间推理导致准确率下降27-56个百分点,缩放70%也显著降效
  • 可诊断模型缺陷,适合评估GUI模型可靠性

GUI grounding模型在标准基准上准确率超85%,但在需要空间推理而非直接元素命名的任务中,准确率下降27至56个百分点。现有基准因仅对每张截图评估一次固定指令而忽略此问题。我们提出GUI-Perturbed,一种受控扰动框架,独立变化视觉场景与指令以测量接地鲁棒性。评估三个同架构的7B模型发现:关系型指令引发所有模型系统性准确率崩溃;浏览器缩放70%导致统计显著性能下降;使用增强数据进行rank-8 LoRA微调反而降低表现。通过沿独立维度扰动,GUI-Perturbed能分离出具体受影响的能力轴——空间推理、视觉鲁棒性、推理校准——提供聚合基准无法获得的诊断信号。我们开源数据集、增强管道及一个微调模型。

原文摘要 · Abstract (English)

GUI grounding models report over 85% accuracy on standard benchmarks, yet drop 27-56 percentage points when instructions require spatial reasoning rather than direct element naming. Current benchmarks miss this because they evaluate each screenshot once with a single fixed instruction. We introduce GUI-Perturbed, a controlled perturbation framework that independently varies visual scenes and instructions to measure grounding robustness. Evaluating three 7B models from the same architecture lineage, we find that relational instructions cause systematic accuracy collapse across all models, a 70% browser zoom produces statistically significant degradation, and rank-8 LoRA fine-tuning with augmented data degrades performance rather than improving it. By perturbing along independent axes, GUI-Perturbed isolates which specific capability axes are affected-spatial reasoning, visual robustness, reasoning calibration-providing diagnostic signal that aggregate benchmarks cannot. We release the dataset, augmentation pipeline, and a fine-tuned model.

GUI理解鲁棒性评估空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。