arXiv:2605.06761cs.AIcs.CV2026-05被引 2

构建可复现的海量网页训练环境,提升视觉网页智能体性能。

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents

论文配图:Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
图 1 · 摘自论文原文
  • 通过HTTP缓存与LLM合成真实网页,实现可复现的多样化训练环境。
  • 在数千个环境中训练,8B模型在多个导航任务上超越同类开源模型。
  • 适合研究网页智能体、强化学习和自动化测试的开发者与研究人员。

网络环境复杂多变,难以规模化生成视觉网页智能体的训练数据。现有方法仅依赖离线轨迹或少数模拟环境,无法覆盖网络多样性。我们提出Weblica(Web Replica)框架,通过两方面实现可复现且可扩展的网页训练环境:1)基于HTTP层缓存,捕获并重放稳定视觉状态,同时保持交互行为;2)利用大语言模型基于真实网站和核心网页导航技能生成环境。借助该框架,我们将强化学习训练扩展至数千个多样环境与任务。最佳模型Weblica-8B在多个网页导航基准上优于同规模开源基线模型,推理步数更少,测试时计算资源增加时表现持续提升,性能媲美API级模型。

原文摘要 · Abstract (English)

The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories for supervised fine-tuning or a handful of simulated environments for RL training, thus failing to capture web diversity. We propose Weblica (Web Replica), a framework for constructing reproducible and scalable web environments. Our framework leverages 1) HTTP-level caching to capture and replay stable visual states while preserving interactive behavior and 2) LLM-based environment synthesis grounded in real-world websites and core web navigation skills. Using this framework, we scale RL training to thousands of diverse environments and tasks. Our best model, Weblica-8B, outperforms open-weight baselines of similar size across multiple web navigation benchmarks while using fewer inference steps, scales favorably with additional test-time compute, and is competitive with API models.

网页智能体强化学习可复现训练大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。