为真实安卓应用设计可验证的评估基准,让手机智能体任务执行更可信。
AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications

- 基于可观测操作规则构建三层评估体系,实现无内部状态下的自动验证。
- 在94个真实应用上完成350项日常任务,人类与模型评估一致率达87.37%。
- 适合研究移动智能体、人机交互及真实场景评测的学者与开发者参考。
GUI基础模型和移动GUI智能体快速发展,但多数评估基准依赖模拟环境或开源应用,难以覆盖真实封闭源码应用。其核心难点在于封闭应用不暴露内部状态,传统自动验证失效。为此,我们提出AndroidDaily,一个涵盖94个高频使用安卓应用(如交通、购物、社交、内容创作等)的大型基准,包含350个真实日常任务。为实现此类不可见环境下的自动可验证评估,我们设计GRADE(Guideline-grounded Reviewer for Automatic Diagnostic Evaluation),一种基于三层次可观测外部规则(操作义务、输出质量、负向约束)的过程感知评估器。GRADE通过追踪智能体的视觉轨迹进行逐步诊断判断,将长周期开放式移动交互转化为可验证评估,无需依赖隐藏内部状态。实验表明,GRADE与人工评估者达成87.37%的一致性。当前最强模型在AndroidDaily上仅达62.0%成功率,揭示出现有推理能力与真实执行之间的显著差距。
原文摘要 · Abstract (English)
The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments or open-source applications, leaving real-world closed-source applications largely unevaluated. The core difficulty is that closed-source applications do not expose internal states, making traditional automatic verification inapplicable. To bridge this gap, we introduce AndroidDaily, a large-scale benchmark comprising 350 realistic daily-use tasks across 94 high-frequency Android applications spanning transportation, shopping, local services, entertainment, content creation, social media, and everyday utilities. To enable automatic and verifiable assessment in these opaque environments, we propose Guideline-grounded Reviewer for Automatic Diagnostic Evaluation (GRADE), a process-aware evaluator built on a three-tiered system of observable external guidelines: operational obligations, output quality, and negative constraints. GRADE tracks the agent's visual trajectory against these criteria and produces step-level diagnostic judgments, turning long-horizon, open-ended mobile interactions into verifiable evaluation without relying on hidden internal states. Experiments show that GRADE achieves 87.37\% agreement with human evaluators. The strongest model reaches a 62.0\% success rate on AndroidDaily, highlighting a substantial gap between current reasoning capabilities and practical execution in realistic mobile workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。