arXiv:2504.11543cs.AI2025-04NeurIPS被引 43

REAL为真实网站模拟提供可复现的智能体评测基准。

REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

  • 构建11个高保真真实网站仿真环境,支持多轮交互
  • 前沿大模型在复杂任务上最高仅41%成功率,暴露能力短板
  • 兼容开源与闭源系统,适合研究智能体可靠性与泛化性

我们提出REAL,一个用于多轮智能体评估的基准与框架,基于11个涵盖电商、旅游、通信和职业社交等领域的高保真、确定性网站仿真。该基准包含112项贴近日常的复杂任务,需结合精准信息检索与状态变更操作。所有交互在完全受控环境中进行,消除安全风险,确保评估的鲁棒性与可复现性。评估框架结合程序化网站状态检查(针对动作类任务)与基于提示的LLM评分(针对信息检索),支持开箱即用的黑盒命令调用,无需修改即可测试各类智能体系统。实证结果表明,前沿语言模型在REAL上的成功率最高仅为41%,凸显其在自主网页导航与任务完成上的显著差距。该框架支持新任务快速集成、可复现评估与规模化后训练数据生成,对推动智能体能力评测与提升具有重要意义。

原文摘要 · Abstract (English)

We introduce REAL, a benchmark and framework for multi-turn agent evaluations on deterministic simulations of real-world websites. REAL comprises high-fidelity, deterministic replicas of 11 widely-used websites across domains such as e-commerce, travel, communication, and professional networking. We also release a benchmark consisting of 112 practical tasks that mirror everyday complex user interactions requiring both accurate information retrieval and state-changing actions. All interactions occur within this fully controlled setting, eliminating safety risks and enabling robust, reproducible evaluation of agent capability and reliability. Our novel evaluation framework combines programmatic checks of website state for action-based tasks with rubric-guided LLM-based judgments for information retrieval. The framework supports both open-source and proprietary agent systems through a flexible evaluation harness that accommodates black-box commands within browser environments, allowing research labs to test agentic systems without modification. Our empirical results show that frontier language models achieve at most a 41% success rate on REAL, highlighting critical gaps in autonomous web navigation and task completion capabilities. Our framework supports easy integration of new tasks, reproducible evaluation, and scalable post-training data generation, marking a significant step forward in evaluating and advancing agent capabilities.

智能体评测网页自动化基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。