arXiv:2502.11357cs.AIcs.HC2025-02ACL被引 74

构建超大规模网页轨迹数据集,推动多模态网页智能体发展

Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents

  • 通过大规模探索与优化生成多样化网页操作轨迹
  • 建成含94,000条成功轨迹、覆盖4.9万唯一网址的数据集
  • 平均每条轨迹成本仅28美分,适合广泛研究者使用

大型多模态模型(LMMs)的进展催生了能自主完成复杂网络任务的智能体。尽管开源LMM智能体在离线基准测试中取得显著进步,但在更真实的在线环境中仍远未达到人类水平。核心瓶颈在于缺乏跨领域、多样且大规模的轨迹级数据集,而人工收集成本高昂。本文提出一种可扩展的方法,合成迄今最大最多样化的轨迹级数据集,包含超过94,000条成功多模态网页轨迹,覆盖49,000个唯一网址、720,000张截图和3300万网页元素。通过广泛网页探索与迭代优化获取多样化任务意图,平均成本仅为28美分/条成功轨迹,具备普惠性。基于该数据集训练出Explorer多模态网页智能体,在Mind2Web-Live、Multimodal-Mind2Web和MiniWob++等离线与在线基准上表现优异。实验表明,数据规模是提升网页智能体能力的关键驱动力。本研究旨在使大规模LMM智能体研究更具可及性。

原文摘要 · Abstract (English)

Recent success in large multimodal models (LMMs) has sparked promising applications of agents capable of autonomously completing complex web tasks. While open-source LMM agents have made significant advances in offline evaluation benchmarks, their performance still falls substantially short of human-level capabilities in more realistic online settings. A key bottleneck is the lack of diverse and large-scale trajectory-level datasets across various domains, which are expensive to collect. In this paper, we address this challenge by developing a scalable recipe to synthesize the largest and most diverse trajectory-level dataset to date, containing over 94K successful multimodal web trajectories, spanning 49K unique URLs, 720K screenshots, and 33M web elements. In particular, we leverage extensive web exploration and refinement to obtain diverse task intents. The average cost is 28 cents per successful trajectory, making it affordable to a wide range of users in the community. Leveraging this dataset, we train Explorer, a multimodal web agent, and demonstrate strong performance on both offline and online web agent benchmarks such as Mind2Web-Live, Multimodal-Mind2Web, and MiniWob++. Additionally, our experiments highlight data scaling as a key driver for improving web agent capabilities. We hope this study makes state-of-the-art LMM-based agent research at a larger scale more accessible.

网页智能体数据合成多模态大规模训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。