用10个网站合成数据训练出能媲美人类标注的网页智能体。
WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web Agents
- 全自动闭环强化学习,把大模型知识压缩成可执行动作。
- 仅用10个网站合成数据,性能相当于人类标注数据在更大环境下的表现。
- 揭示不同大模型的具身潜力,为模型评估提供新维度。
当前GUI智能体训练受限于不安全的实时网络交互或昂贵稀缺的人工标注数据与环境。本文指出,相较于数据量,更关键的是将大语言模型(LLM)的隐含知识高效压缩为可执行行为。我们提出WebFactory,一种完全自动化的闭环强化学习流水线,系统性地将LLM编码的互联网智能转化为高效、具身的行动。该流程包括可扩展环境合成、知识感知任务生成、LLM驱动轨迹收集、分解奖励强化学习训练及系统化评估。令人瞩目的是,仅在WebFactory中10个网站的合成数据上训练的智能体,其性能可媲美在更大规模人工标注环境中训练的同类智能体。该优势在内部离线与在线迁移基准测试中一致体现,且显著优于基础大模型。我们进一步揭示了不同LLM基础模型的“具身潜力”,提供了新的模型评估维度。本工作构建了一种可扩展、低成本的范式,将被动的互联网知识转化为主动的具身智能,是迈向通用交互智能体的关键一步。
原文摘要 · Abstract (English)
Current paradigms for training GUI agents are fundamentally limited by a reliance on either unsafe, non-reproducible live web interactions or costly, scarce human-crafted data and environments. We argue this focus on data volume overlooks a more critical factor: the efficiency of compressing a large language model's (LLM) latent knowledge into actionable agent behavior. We introduce WebFactory, a novel, fully automated closed-loop reinforcement learning pipeline for GUI agents, systematically compressing LLM-encoded internet intelligence into efficient, grounded actions. Our pipeline features a process of scalable environment synthesis, knowledge-aware task generation, LLM-powered trajectory collection, decomposed reward RL training, and systematic agent evaluation. Remarkably, our agent demonstrates exceptional data efficiency and generalization. Trained on synthetic data from only 10 websites within WebFactory, it achieves performance comparable to GUI agents trained on the same amount of human-annotated data from a much larger set of environments. This superior performance is consistent across our internal offline and online transfer benchmarks, where our agent also significantly outperforms the base foundation model. We further provide critical insights into the "embodiment potential" of different LLM foundations, offering a new axis for model evaluation. This work presents a scalable and cost-effective paradigm for transforming passive internet knowledge into active, grounded intelligence, marking a critical step towards general-purpose interactive agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。