构建跨平台GUI自动化评估框架,揭示高效自动化关键在于精准定位与合理规划。
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
- 分四层评估跨平台GUI代理能力:内容理解、元素定位、任务执行与协作。
- 发现精准视觉定位是任务成功的核心,模型普遍存在冗余操作,效率低下。
- 适合研究自动化代理、人机交互及多平台智能体的开发者和研究人员。
我们提出MMBench-GUI,一个针对Windows、macOS、Linux、iOS、Android和Web平台的分层评估框架,用于评测GUI自动化代理。该框架包含四个层级:GUI内容理解、元素定位、任务自动化与任务协作,覆盖了GUI代理所需的核心能力。此外,我们引入新的效率-质量面积(EQA)指标,评估在线自动化场景中的执行效率。通过MMBench-GUI,我们发现精准视觉定位是任务成功率的关键决定因素,强调集成专用定位模块的模块化框架带来的显著优势。同时,可靠GUI自动化还需强大的任务规划与跨平台泛化能力,长上下文记忆、广阔动作空间和长期推理至关重要。更重要的是,任务效率这一维度被严重忽视,所有模型均存在明显低效问题,即使任务完成也常伴随大量冗余步骤。精准定位、有效规划与早期终止策略的结合,是实现真正高效可扩展自动化所不可或缺的。基准代码、评估数据与运行环境将公开发布于https://github.com/open-compass/MMBench-GUI。
原文摘要 · Abstract (English)
We introduce MMBench-GUI, a hierarchical benchmark for evaluating GUI automation agents across Windows, macOS, Linux, iOS, Android, and Web platforms. It comprises four levels: GUI Content Understanding, Element Grounding, Task Automation, and Task Collaboration, covering essential skills for GUI agents. In addition, we propose a novel Efficiency-Quality Area (EQA) metric to assess GUI agent execution efficiency in online automation scenarios. Through MMBench-GUI, we identify accurate visual grounding as a critical determinant of overall task success, emphasizing the substantial benefits of modular frameworks that integrate specialized grounding modules. Furthermore, to achieve reliable GUI automation, an agent requires strong task planning and cross-platform generalization abilities, with long-context memory, a broad action space, and long-term reasoning playing a critical role. More important, task efficiency remains a critically underexplored dimension, and all models suffer from substantial inefficiencies, with excessive redundant steps even when tasks are ultimately completed. The integration of precise localization, effective planning, and early stopping strategies is indispensable to enable truly efficient and scalable GUI automation. Our benchmark code, evaluation data, and running environment will be publicly available at https://github.com/open-compass/MMBench-GUI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。