arXiv:2503.02403cs.AI2025-03被引 9

自动评估移动智能体,无需人工干预。

AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents

  • 用界面状态变化自动生成任务奖励信号
  • 自主评估准确率达94%,接近人工水平
  • 可快速评测先进智能体,发现性能瓶颈

对移动智能体的全面评估能显著推动其发展与实际应用。然而,现有基准因需大量人工定义任务奖励信号和编写评估代码,缺乏实用性和可扩展性。我们提出AutoEval,一种无需人工干预即可测试移动智能体的评估框架。该方法设计了界面状态变化表示,用于自动生成任务奖励信号,并引入判断系统实现自主评估。实验表明,AutoEval生成的奖励信号与人工标注信号高度相关,自主评估准确率最高达94%,与人工评估相当。最后,我们使用该框架评估了当前最先进的移动智能体,揭示了其性能表现与局限性。

原文摘要 · Abstract (English)

Comprehensive evaluation of mobile agents can significantly advance their development and real-world applicability. However, existing benchmarks lack practicality and scalability due to the extensive manual effort in defining task reward signals and implementing evaluation codes. We propose AutoEval, an evaluation framework which tests mobile agents without any manual effort. Our approach designs a UI state change representation which can be used to automatically generate task reward signals, and employs a Judge System for autonomous evaluation. Evaluation shows AutoEval can automatically generate reward signals with high correlation to human-annotated signals, and achieve high accuracy (up to 94%) in autonomous evaluation comparable to human evaluation. Finally, we evaluate state-of-the-art mobile agents using our framework, providing insights into their performance and limitations.

智能体评估自动化测试移动智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。