arXiv:2604.24964cs.LGcs.CL2026-04被引 13

构建真实长时网页任务基准,评估模型跨网站持续操作能力。

Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

论文配图:Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
图 1 · 摘自论文原文
  • 设计200个来自真实浏览会话的长周期多站点任务
  • 44.5%成功率揭示当前模型仍有巨大提升空间
  • 引入评分细则与效率指标,更精准衡量复杂任务表现

现有网页智能体评测集中于短时、单站点任务,前沿模型已接近饱和。但真实网络使用涉及长时间、跨站点的复杂流程,如跨平台比价、多服务订票或整合多查询信息。为此,我们提出Odysseys:一个基于真实浏览会话的200个长周期网页任务基准,在实时互联网上进行评估。发现二元成败评价不适用于长周期场景,引入平均6.1项分级评分细则,显著提升与人类判断的一致性,并优于常用大模型作为裁判的轨迹级评估。测试多个领先模型后发现,最强模型成功率达44.5%,仍有大幅提升空间。此外,强调效率为首要考量,提出轨迹效率指标(每步得分),发现顶级模型仅达1.15%,表明亟需高效且可持续的任务执行能力。Odysseys为开放网络中长周期智能体能力评估提供真实基准,推动向可长时间协作的计算机助手迈进。数据与代码已开源。

原文摘要 · Abstract (English)

Existing web agent benchmarks have largely converged on short, single-site tasks that frontier models are approaching saturation on. However, real world web use consists of long-horizon, multi-site workflows. Common web navigation tasks, such as comparing products across different domains, planning trips across multiple services, or summarizing information from multiple search queries, require sustained context and cross-site reasoning over potentially hours of browsing. To capture and evaluate such behaviors, we introduce Odysseys: a benchmark of 200 long-horizon web tasks derived from real world browsing sessions evaluated on the live Internet. We find that binary pass/fail evaluation is inadequate for long-horizon settings and introduce a rubric-based evaluation, annotating each Odysseys task with an average of 6.1 graded rubrics. We demonstrate that this yields higher agreement with humans and provides a more fine-grained signal than commonly used trajectory-level LLM-as-a-judge evaluation metrics. We tested several leading frontier models and find that the strongest models achieve a success rate of 44.5%, which leaves substantial room for future improvements. Beyond task success, we argue that efficiency is a first-class concern for long-horizon agents. We introduce a Trajectory Efficiency metric (rubric score per step) and find that even frontier agents achieve only 1.15%, marking an evident need for agents that can succeed efficiently and not simply eventually. Odysseys isolates the critical evaluation of long-horizon proficiency in open-web environments, providing a realistic benchmark to measure progress towards computer-use agents that can potentially productively operate for hours. We release our tasks, evaluation scripts, and other results at https://odysseys-website.pages.dev

网页代理长时任务基准评测效率评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。