arXiv:2602.12544cs.AI2026-02被引 4

自动生成网页代理训练数据,提升任务完成度评估精度。

Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation

  • 用约束式框架细粒度评估任务进展,利用部分成功轨迹
  • 在20个网站的预订任务上,小模型性能媲美或超越商用系统
  • 适合做网页自动化、智能代理开发的研究者与工程师

我们提出一种可扩展的自动生成高质量网页代理训练数据的流程。核心挑战在于轨迹评估——量化任务完成进度。为此,我们引入一种新颖的基于约束的评估框架,实现对任务进展的细粒度判断,从而能有效利用部分成功的轨迹,显著扩大可用训练数据规模。我们在新提出的基准 BookingArena 上进行评估,该基准涵盖20个热门网站上的复杂预订任务。结果表明,我们提炼的轻量级学生模型性能优于开源方法,并达到或超过商业系统水平,同时模型规模更小。本工作解决了高效构建多样化、真实网页交互数据集的难题,为复杂结构化网页任务提供了系统化的评估方法。

原文摘要 · Abstract (English)

We present a scalable pipeline for automatically generating high-quality training data for web agents. In particular, a major challenge in identifying high-quality training instances is trajectory evaluation - quantifying how much progress was made towards task completion. We introduce a novel constraint-based evaluation framework that provides fine-grained assessment of progress towards task completion. This enables us to leverage partially successful trajectories, which significantly expands the amount of usable training data. We evaluate our method on a new benchmark we propose called BookingArena, which consists of complex booking tasks across 20 popular websites, and demonstrate that our distilled student model outperforms open-source approaches and matches or exceeds commercial systems, while being a significantly smaller model. Our work addresses the challenge of efficiently creating diverse, realistic web interaction datasets and provides a systematic evaluation methodology for complex structured web tasks.

网页代理自动标注评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。