新基准测试让AI理解界面操作背后的逻辑,而非仅识别按钮。
AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark

- 用视觉语言模型与人工协作,从多平台截图生成层级功能区。
- 覆盖6个系统2753个任务,测试界面语义、定位与状态预测能力。
- 发现开源模型懂功能,商用模型会描述,但都难搞复杂操作逻辑。
能够导航图形用户界面(GUI)的自主智能体有望彻底改变数字生产力。然而,真正的数字自治不仅依赖于反应式元素匹配,更需要对界面动态建立预测性心智模型,并预判交互后的“数字世界状态”。尽管现代视觉-语言模型(VLMs)具备感知能力,现有基准仍分裂(聚焦黑箱任务完成或静态浅层定位),无法评估智能体是否真正理解界面的隐含功能与状态转换逻辑。为此,我们提出AutoGUI-v2,一个全面的基准,用于评估深度GUI功能理解与交互结果预测能力。通过创新的VLM-人类协同流水线,递归解析多平台截图,构建出层次化功能区域,生成多样化评估任务。该基准涵盖2753个任务,覆盖六个操作系统,在区域与元素级语义、定位及动态状态预测方面严格测试智能体表现。评估显示VLMs存在显著差异:微调过代理数据的开源模型(如Qwen3-VL)在功能定位上表现优异,而商用模型(如Gemini-2.5-Pro-Thinking)则在功能描述上占优。关键的是,所有模型在不常见操作的复杂交互逻辑上均表现不佳,表明深层功能理解仍是重大挑战。AutoGUI-v2系统性测量这些基础能力,为下一代GUI智能体的发展提供了新视角。
原文摘要 · Abstract (English)
Autonomous agents capable of navigating Graphical User Interfaces (GUIs) hold the potential to revolutionize digital productivity. However, achieving true digital autonomy extends beyond reactive element matching; it necessitates a predictive mental model of interface dynamics and the ability to foresee the "digital world state" resulting from interactions. Despite the perceptual capabilities of modern Vision-Language Models (VLMs), existing benchmarks remain bifurcated (focusing either on black-box task completion or static, shallow grounding), thereby failing to assess whether agents truly comprehend the implicit functionality and transition logic of GUIs. To bridge this gap, we introduce AutoGUI-v2, a comprehensive benchmark designed to evaluate deep GUI functionality understanding and interaction outcome prediction. We construct the benchmark using a novel VLM-human collaborative pipeline that recursively parses multi-platform screenshots into hierarchical functional regions to generate diverse evaluation tasks. Providing 2,753 tasks across six operating systems, AutoGUI-v2 rigorously tests agents on region and element-level semantics, grounding, and dynamic state prediction. Our evaluation reveals a striking dichotomy in VLMs: while open-source models fine-tuned on agent data (e.g., Qwen3-VL) excel at functional grounding, commercial models (e.g., Gemini-2.5-Pro-Thinking) dominate in functionality captioning. Crucially, all models struggle with complex interaction logic of uncommon actions, highlighting that deep functional understanding remains a significant hurdle. By systematically measuring these foundational capabilities, AutoGUI-v2 offers a new lens for advancing the next generation of GUI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。