arXiv:2605.29218cs.AIcs.CL2026-05ACL

构建可扩展的长时序网页任务生成框架,解决智能体训练数据稀缺问题。

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

论文配图:GTA: Generating Long-Horizon Tasks for Web Agents at Scale
图 1 · 摘自论文原文
  • 通过爬取+检索种子+上下文生成,实现高效自动任务创建
  • 在50多个网站上生成多语言多跳任务,支持动态评估
  • 提供可复现的基准测试,揭示人类与智能体性能差距

网页智能体结合大语言模型与浏览、工具使用能力,有望成为开放网络助手。然而,进展受限于缺乏可扩展的过程级监督。现有基准多为人工构建,仅提供起点-目标标注,缺少中间轨迹;近期自动化生成方法成本高、有偏且浅显。这些限制阻碍了需泛化至真实多跳跨页任务的智能体训练与评估。本文提出可扩展框架GTA,整合爬取、基于检索的种子生成、上下文生成与自动化质量控制,生成配对可执行轨迹的真实任务。该设计将爬取与生成解耦以提升效率,基于站点图结构确保任务组合性,并通过确定性重放与系统验证保证密集监督。我们在超过50个涵盖电商、政府、论坛和新闻的网站上实例化该流程,实现多语言、多跳覆盖。生成的基准揭示显著的人机性能差距,并支持细致诊断。贡献包括:(i) 形式化多跳网页智能体任务生成,(ii) 提出高效且经验证的自动数据生成流水线,(iii) 发布可复现的动态基准与评估体系。

原文摘要 · Abstract (English)

Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limited by the lack of scalable, process-level supervision. Existing benchmarks are largely manually constructed, providing only coarse start-goal annotations without intermediate trajectories, while recent automatic generation efforts remain expensive, biased, and shallow. These limitations prevent reliable training and evaluation of agents that must generalize to realistic, multi-hop, cross-page tasks. We introduce a scalable framework, GTA, that integrates crawling, retrieval-based seeding, in-context generation, and automated quality control to produce realistic tasks paired with executable trajectories. This design decouples crawling from generation for greater efficiency, grounds tasks in the site graph to enforce compositionality, and ensures dense supervision through deterministic replays and systematic validation. We instantiate the pipeline on over 50 websites covering e-commerce, government, forums, and news, with multilingual and multi-hop coverage. The resulting benchmark reveals a significant human-agent performance gap and enables detailed diagnostics. Our contributions are three-fold: (i) formalizing multi-hop web-agent task generation, (ii) proposing an efficient and validated pipeline for automatic data creation, and (iii) releasing a dynamic benchmark with reproducible evaluation.

网页智能体任务生成多跳推理自动标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。