构建桌面环境GUI定位基准,测试大模型在复杂窗口下的鲁棒性。
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments

- 通过参数化生成多窗口遮挡场景,模拟真实工作流分布
- 5个顶级MLLM在部分遮挡下准确率下降,暴露可靠性短板
- 适合研究GUI自动化、多模态模型鲁棒性的学者使用
多模态大语言模型(MLLMs)虽已革新GUI自动化,但其有效性主要基于理想化的单层界面。本文指出一个关键可靠性缺口:当前先进智能体在多窗口堆叠、遮挡和视觉杂乱的真实桌面环境中面临显著鲁棒性挑战。为此,我们提出WinDeskGround——一个专为评估GUI定位鲁棒性而设计的新基准与合成框架。不同于静态数据集,该框架通过控制窗口遮挡程度、布局密度和语义相似性,参数化生成复杂桌面场景,从而模拟真实工作流的分布偏移。我们构建了一个包含1,356对高保真指令-目标样本的多样化元数据集,并对五个领先MLLM进行了全面评估。结果表明,尽管顶尖代理在简化环境下表现优异,但在部分遮挡条件下准确率明显下降。WinDeskGround为评估和推进真实环境中的GUI智能体鲁棒性提供了宝贵基准。代码已开源:https://github.com/ZZZhr-1/WinDeskGround。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interfaces. This paper identifies a critical reliability gap: state-of-the-art agents face distinct robustness challenges in real-world desktop environments characterized by multi-window stacking, occlusion, and visual clutter. To address this, we introduce WinDeskGround, a novel benchmark and synthesis framework tailored for evaluating GUI grounding robustness. Unlike static datasets, our framework parametrically generates complex desktop scenarios by controlling window occlusion, layout density, and semantic similarity, thereby simulating the distribution shifts of authentic workflows. We construct a diverse meta-dataset of 1,356 high-fidelity instruction-target pairs and conduct comprehensive evaluations of five leading MLLMs. Our results demonstrate that while top-tier agents excel in simplified settings, their accuracy declines under partial occlusion. WinDeskGround provides a valuable benchmark to facilitate the assessment and advancement of GUI agent robustness in realistic environments. The code is available at https://github.com/ZZZhr-1/WinDeskGround.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。