arXiv:2604.06182cs.HCcs.AI2026-04被引 3

打造更贴近真实用户使用的手机界面智能体评测基准。

VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics

论文配图:VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
图 1 · 摘自论文原文
  • 以用户意图驱动设计任务,还原真实使用场景。
  • 揭示现有智能体在感知与记忆上存在明显短板。
  • 适合评估智能体在复杂环境下的鲁棒性与实用性。

现有手机界面智能体的在线评测基准多聚焦特定应用且任务同质化,无法反映真实移动使用中的多样性与不稳定性。为此,我们提出VenusBench-Mobile,一个面向通用手机界面智能体的挑战性在线评测基准,强调真实、用户中心的评估条件。该基准构建两大核心支柱:通过用户意图驱动的任务设计明确评估内容,通过能力导向的标注体系实现对智能体行为的细粒度分析。对先进移动端智能体的广泛评估显示,其性能相较于以往基准存在显著差距,表明VenusBench-Mobile提出了更具挑战性和现实意义的任务;当前智能体距离实际部署仍有很大距离。诊断分析进一步表明,失败主要源于感知与记忆能力不足,这些缺陷在粗粒度评估中被掩盖。即使最强智能体在环境变化下成功率也接近零,凸显其在真实场景中的脆弱性。基于这些发现,我们认为VenusBench-Mobile为推动手机界面智能体在真实世界中可靠部署提供了重要基础。代码与数据已开源于https://github.com/inclusionAI/UI-Venus/tree/VenusBench-Mobile。

原文摘要 · Abstract (English)

Existing online benchmarks for mobile GUI agents remain largely app-centric and task-homogeneous, failing to reflect the diversity and instability of real-world mobile usage. To this end, we introduce VenusBench-Mobile, a challenging online benchmark for evaluating general-purpose mobile GUI agents under realistic, user-centric conditions. VenusBench-Mobile builds two core evaluation pillars: defining what to evaluate via user-intent-driven task design that reflects real mobile usage, and how to evaluate through a capability-oriented annotation scheme for fine-grained agent behavior analysis. Extensive evaluation of state-of-the-art mobile GUI agents reveals large performance gaps relative to prior benchmarks, indicating that VenusBench-Mobile poses substantially more challenging and realistic tasks and that current agents remain far from reliable real-world deployment. Diagnostic analysis further shows that failures are dominated by deficiencies in perception and memory, which are largely obscured by coarse-grained evaluations. Moreover, even the strongest agents exhibit near-zero success under environment variations, highlighting their brittleness in realistic settings. Based on these insights, we believe VenusBench-Mobile provides an important stepping stone toward robust real-world deployment of mobile GUI agents. Code and data are available at https://github.com/inclusionAI/UI-Venus/tree/VenusBench-Mobile.

手机智能体评测基准用户中心

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。