首个专用于评估自动探索用户界面的基准,推动智能体高效操作界面。
Toward Autonomous UI Exploration: The UIExplorer Benchmark
- 构建了结构化与屏幕两种模式下的标准化界面探索评测环境。
- 提出新指标hUFO,量化探索效率,领先模型达人类77.2%性能。
- 适合研究自动化界面操作、智能体交互与测试数据生成的学者。
自主智能体需掌握用户界面(UI)探索能力以可靠完成任务,但该关键环节缺乏系统性评估。本文提出首个专注于界面探索的基准UIExplore-Bench,涵盖结构化模式(可访问布局信息如DOM树)和屏幕模式(仅依赖截图与鼠标键盘输入),在标准GitLab沙箱环境中分三个难度层级进行评估。将探索定义为最大化可操作组件发现数量,并引入人类归一化界面功能观测率(hUFO)作为量化指标。实验表明,UIExplore-AlGo在2000步内取得领先表现,结构化模式下最高达人类77.2%,屏幕模式下达59.0%,尤其在稀疏场景表现突出。结果揭示当前智能体与人类专家一小时探索水平间存在显著差距,显示巨大改进空间。我们公开发布基准环境、探索数据集及评估工具套件,以推动高效界面探索策略及其下游应用(如经验驱动的任务完成与自动化训练数据生成)的研究。
原文摘要 · Abstract (English)
Autonomous agents must know how to explore user interfaces (UIs) for reliable task solving, yet systematic evaluation of this crucial phase is lacking. We introduce UIExplore-Bench, the first benchmark explicitly dedicated to UI exploration. The benchmark evaluates agents with either Structured mode (granting access to layout information like DOM trees) or Screen mode (relying on GUI-only observations such as screenshots and human-like mouse/keyboard interactions) across three levels in a standardized GitLab sandbox environment. We formalize exploration as the process of maximizing the set of actionable UI components discovered and propose a metric, human-normalized UI-Functionalities Observed (hUFO), to quantify the effectiveness of exploration. Our results show that UIExplore-AlGo achieves the leading mean hUFO scores, reaching up to 77.2% of human performance in Structured mode and 59.0% in Screen mode at 2,000 steps, particularly excelling at the Sparse level. The results highlight the relevance of our benchmark, as current agents show a substantial performance gap compared to one hour of human expert exploration, indicating ample room for future advancements. We publicly release the benchmark environment, an exploration dataset, and an evaluation suite to catalyze research into efficient UI exploration strategies and their downstream applications, such as experience-driven task completion and automated training data generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。