arXiv:2503.03056cs.LG2025-03被引 2

构建真实世界自主智能体评测基准,覆盖芯片布局、网页导航等场景。

A2Perf: Real-World Autonomous Agents Benchmark

  • 设计三个贴近现实的测试环境:芯片布局、网页导航、四足机器人运动。
  • 量化评估任务性能、泛化能力、资源效率与可靠性,揭示算法差异。
  • 适合研究真实世界智能体的学者,尤其关注落地应用与成本效益者。

自主智能体与系统广泛应用于机器人、数字助手、组合优化等领域,面临通用性、可靠性与资源效率等共性挑战。现有方法如强化学习与模仿学习各有优劣,但缺乏统一的评测基准来衡量其在真实问题上的表现。本文提出A2Perf——一个包含芯片布局、网页导航、四足机器人运动三个真实场景的基准。该基准提供任务性能、泛化能力、系统资源效率与可靠性四项核心指标,并验证了网页导航智能体在消费级硬件上可实现接近人类反应时间的延迟,揭示了四足运动算法间的可靠性权衡,并量化了不同学习方法在芯片设计中的能耗成本。此外,引入数据成本度量以评估离线数据获取代价,支持模仿学习与混合方法的公平比较。基准包含多个标准基线,支持方法间直接对比。A2Perf为开源项目,旨在长期服务研究社区。

原文摘要 · Abstract (English)

Autonomous agents and systems cover a number of application areas, from robotics and digital assistants to combinatorial optimization, all sharing common, unresolved research challenges. It is not sufficient for agents to merely solve a given task; they must generalize to out-of-distribution tasks, perform reliably, and use hardware resources efficiently during training and inference, among other requirements. Several methods, such as reinforcement learning and imitation learning, are commonly used to tackle these problems, each with different trade-offs. However, there is a lack of benchmarking suites that define the environments, datasets, and metrics which can be used to provide a meaningful way for the community to compare progress on applying these methods to real-world problems. We introduce A2Perf--a benchmark with three environments that closely resemble real-world domains: computer chip floorplanning, web navigation, and quadruped locomotion. A2Perf provides metrics that track task performance, generalization, system resource efficiency, and reliability, which are all critical to real-world applications. Using A2Perf, we demonstrate that web navigation agents can achieve latencies comparable to human reaction times on consumer hardware, reveal reliability trade-offs between algorithms for quadruped locomotion, and quantify the energy costs of different learning approaches for computer chip-design. In addition, we propose a data cost metric to account for the cost incurred acquiring offline data for imitation learning and hybrid algorithms, which allows us to better compare these approaches. A2Perf also contains several standard baselines, enabling apples-to-apples comparisons across methods and facilitating progress in real-world autonomy. As an open-source benchmark, A2Perf is designed to remain accessible, up-to-date, and useful to the research community over the long term.

自主智能体评测基准真实世界资源效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。