用间接知识生成海量模拟操作数据,让数字代理更智能高效
Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale
- 将在线教程等间接知识转化为可直接训练的模拟操作数据
- 用10万条合成数据微调7B模型,在多个网页任务上超越现有模型
- 合成数据成本仅为人工数据的3%,且效果更优
大语言模型如今可作为自主代理与数字环境交互并完成特定目标(如安排在线会议),但准确率仍不理想,主要因缺乏大规模直接示范数据。人工获取标注数据成本高,而通过探索或强化学习自动收集又依赖复杂的环境和内容配置,导致数据集覆盖场景有限。另一方面,存在大量可间接辅助任务完成的知识,如面向人类的在线教程。本文提出Synatra,一种将此类间接知识规模化转化为直接监督信号的方法。我们定义了不同类型的间接知识,系统研究其来源、直接示范结构编码方式及转化方法。利用10万条合成示范数据微调7B CodeLlama模型,结果表明该代理在Mind2Web、MiniWoB++和WebArena三个基于网页的任务基准上均优于同类规模模型,并在WebArena和Mind2Web上超越GPT-3.5。此外,合成示范数据成本仅为人工数据的3%(每条$0.031),且在有限领域内收集的相同数量人工数据对比下,合成数据表现更优。
原文摘要 · Abstract (English)
LLMs can now act as autonomous agents that interact with digital environments and complete specific objectives (e.g., arranging an online meeting). However, accuracy is still far from satisfactory, partly due to a lack of large-scale, direct demonstrations for digital tasks. Obtaining supervised data from humans is costly, and automatic data collection through exploration or reinforcement learning relies on complex environmental and content setup, resulting in datasets that lack comprehensive coverage of various scenarios. On the other hand, there is abundant knowledge that may indirectly assist task completion, such as online tutorials that were created for human consumption. In this work, we present Synatra, an approach that effectively transforms this indirect knowledge into direct supervision at scale. We define different types of indirect knowledge, and carefully study the available sources to obtain it, methods to encode the structure of direct demonstrations, and finally methods to transform indirect knowledge into direct demonstrations. We use 100k such synthetically-created demonstrations to finetune a 7B CodeLlama, and demonstrate that the resulting agent surpasses all comparably sized models on three web-based task benchmarks Mind2Web, MiniWoB++ and WebArena, as well as surpassing GPT-3.5 on WebArena and Mind2Web. In addition, while synthetic demonstrations prove to be only 3% the cost of human demonstrations (at $0.031 each), we show that the synthetic demonstrations can be more effective than an identical number of human demonstrations collected from limited domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。