用五万条交互轨迹微调大模型,让AI助手更通用
AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories
- 构建超大规模交互轨迹数据集,覆盖16项任务与5类技能
- 微调后模型在多任务中表现显著提升,验证数据规模有效性
- 适合想训练通用智能体的研究者与开发者
在智能体-环境交互轨迹数据上进行微调,有望激发开源大语言模型的通用智能体能力。本文提出AgentBank,目前规模最大、包含超过5万条多样化高质量交互轨迹的数据集,涵盖16个任务及五大不同智能体技能维度。通过创新的标注流程,有效降低难度偏差。基于该数据集,我们对大语言模型进行微调,得到一系列智能体模型Samoyed。对比实验表明,扩大交互轨迹数据规模能有效获得通用智能体能力。附加研究还揭示了轨迹微调与技能泛化的一些关键发现。
原文摘要 · Abstract (English)
Fine-tuning on agent-environment interaction trajectory data holds significant promise for surfacing generalized agent capabilities in open-source large language models (LLMs). In this work, we introduce AgentBank, by far the largest trajectory tuning data collection featuring more than 50k diverse high-quality interaction trajectories which comprises 16 tasks covering five distinct agent skill dimensions. Leveraging a novel annotation pipeline, we are able to scale the annotated trajectories and generate a trajectory dataset with minimized difficulty bias. Furthermore, we fine-tune LLMs on AgentBank to get a series of agent models, Samoyed. Our comparative experiments demonstrate the effectiveness of scaling the interaction trajectory data to acquire generalized agent capabilities. Additional studies also reveal some key observations regarding trajectory tuning and agent skill generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。