将用户浏览行为转化为可复用的自然语言技能,提升浏览器代理的可扩展性。
Scalable Behaviour Cloning on Browser Using via Skill Distillation

- 通过技能蒸馏将用户操作轨迹转为自然语言技能
- 构建技能图谱实现技能合并而非无限累积
- 适合想构建智能浏览器代理的研究者与开发者
互联网用户通过浏览器完成了从软件开发、文档编辑到搜索、表单填写及企业工作流等大量高阶任务,形成海量可复用的浏览器技能。我们提出,浏览器代理的瓶颈在于不完整信息下的决策,而非底层操作;而这些决策先验已隐含在人类交互轨迹中。为此,本文研究基于技能蒸馏的可扩展行为克隆:将用户交互轨迹转化为紧凑的自然语言技能,使代理可直接阅读、检索、复用和组合。进一步将蒸馏出的技能组织为技能图谱,实现技能增长通过整合而非无限制积累。这表明浏览器代理的可扩展性可能更多来自互联网用户的集体经验,而非人工设计任务。项目地址:https://lab.einsia.ai/browserbc/
原文摘要 · Abstract (English)
Internet users collectively perform an enormous range of skilled work through web browsers, from software development and document editing to search, forms, and enterprise workflows, making human browsing a highly scalable but under-exploited source of reusable browser skills. We argue that the bottleneck for browser agents is decision-making under incomplete information rather than low-level operation, and that the priors agents lack are already implicit in human interaction traces. We therefore study scalable behavior cloning for browser agents via skill distillation, converting user interaction trajectories into compact natural-language skills that agents can read, retrieve, reuse, and compose directly. We further organize the distilled skills into a skill graph so that growth proceeds through consolidation rather than unbounded accumulation. This suggests that the scalability of browser agents may come less from manually designed tasks and more from the collective skills already expressed by internet users. Our project is available at: https://lab.einsia.ai/browserbc/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。