打造能长期稳定操作真实网页的浏览器智能体,解决实际应用中的容错与复杂界面问题。
Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents

- 构建统一框架,从执行到评估全程优化,支持错误恢复与复杂界面交互。
- 在真实任务基准上达成80.6%成功率,刷新开源模型性能纪录。
- 适用于多场景代理任务,展现强大通用智能能力,适合研究长时序决策系统者参考。
浏览器智能体在简短、干净的演示中表现良好,但真实部署面临巨大挑战:需在动态网站上持续做出数十次决策,同时处理错误并应对复杂用户界面。我们主张,缩小这一差距需在执行、监督、优化和评估等全链路实现对齐,而非仅依赖规模扩展。本文提出Wuying-Browser-Agent,一个涵盖各环节的统一框架。结构化浏览器引擎提供稳定的执行原语与面向决策的上下文管理;反射与面向界面的课程微调(RUIC-SFT)显式训练于恢复轨迹与复杂界面交互;基于分歧感知的在线GRPO(DAO-GRPO)通过基于势能的奖励塑形与分歧感知步权重,提升长时程信用分配。此外,我们引入BrowserBench,一个包含350项任务、平均37.9步的双语真实网页基准,以暴露长时程失败模式。Wuying-Browser-Agent-27B在WebVoyager上达80.6%,Online-Mind2Web上66.7%,BrowserBench上65.1%,确立开源新标杆。该框架还可泛化至非浏览器任务,在Tau2-Bench、Claw-Eval与BFCL-v4上平均得分73.8。
原文摘要 · Abstract (English)
Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally different: agents must sustain dozens of decisions on live websites while recovering from mistakes and navigating complex UIs. We argue that closing this gap requires alignment at every level of the pipeline, including execution, supervision, optimization, and evaluation, rather than scale alone. We present Wuying-Browser-Agent, a unified framework that addresses each of these levels. A structured browser harness provides stable execution primitives and decision-oriented context management. Reflection and UI-specialized Curriculum SFT (RUIC-SFT) explicitly trains on recovery trajectories and complex-UI interactions. Divergence-Aware Online GRPO (DAO-GRPO) improves long-horizon credit assignment through potential-based reward shaping and divergence-aware step weighting. Finally, we introduce BrowserBench, a bilingual real-web benchmark of 350 tasks averaging 37.9 steps, because most existing benchmarks are too short to expose long-horizon failure modes. Wuying-Browser-Agent-27B achieves 80.6\% on WebVoyager, 66.7\% on Online-Mind2Web, and 65.1\% on BrowserBench, establishing a new open-source state of the art on browser-use benchmarks. The same pipeline also transfers beyond browser use, demonstrating strong general agentic ability and reaching an average score of 73.8 on Tau2-Bench, Claw-Eval, and BFCL-v4.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。