arXiv:2512.21302cs.CV2025-12被引 2

构建复杂安卓交互评估框架,测试智能代理长期任务能力

AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents

  • 基于真实场景设计571个跨域长流程任务,平均需26步完成
  • 最佳模型仅12.7%成功率,ATP达50.47%,暴露关键挑战
  • 支持静态与动态评估,适配真实环境中的异常与多路径

图形用户界面(GUI)代理可通过自动化移动端高频长延迟任务显著提升效率。然而现有评测基准仍受限于少量应用、简单任务及粗粒度指标。为此,我们提出AndroidLens,一个面向移动GUI代理的挑战性评估框架,包含571个在中英文环境下运行的长延迟任务,每个任务平均需超过26步完成。框架具备三大特性:(1) 来自38个领域的实际用户场景,涵盖多约束、多目标及领域特定等复杂类型任务;(2) 静态评估保留真实世界异常,允许多条有效路径以降低偏差;(3) 动态评估采用里程碑机制,通过平均任务进度(ATP)实现细粒度进展测量。评估显示,即使最优模型也仅达12.7%任务成功率和50.47% ATP。我们还揭示了真实环境中存在的关键挑战,包括环境异常、自适应探索与长期记忆保持。

原文摘要 · Abstract (English)

Graphical user interface (GUI) agents can substantially improve productivity by automating frequently executed long-latency tasks on mobile devices. However, existing evaluation benchmarks are still constrained to limited applications, simple tasks, and coarse-grained metrics. To address this, we introduce AndroidLens, a challenging evaluation framework for mobile GUI agents, comprising 571 long-latency tasks in both Chinese and English environments, each requiring an average of more than 26 steps to complete. The framework features: (1) tasks derived from real-world user scenarios across 38 domains, covering complex types such as multi-constraint, multi-goal, and domain-specific tasks; (2) static evaluation that preserves real-world anomalies and allows multiple valid paths to reduce bias; and (3) dynamic evaluation that employs a milestone-based scheme for fine-grained progress measurement via Average Task Progress (ATP). Our evaluation indicates that even the best models reach only a 12.7% task success rate and 50.47% ATP. We also underscore key challenges in real-world environments, including environmental anomalies, adaptive exploration, and long-term memory retention.

GUI代理长延迟任务安卓评测多步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。