开源大模型xLAM提升智能体工具使用能力,性能超GPT-4
xLAM: A Family of Large Action Models to Empower AI Agent Systems
- 构建统一数据管道,融合增强多源数据训练模型
- 5个模型从1B到8×22B参数,工具调用表现领先
- 适合研究自主智能体与开源模型应用的开发者
由大语言模型驱动的自主智能体受到广泛关注。然而,开源社区在开发专用智能体模型时面临挑战,主要由于高质量智能体数据集稀缺及该领域缺乏标准协议。我们推出并公开发布xLAM系列大动作模型,专为智能体任务设计。该系列包含5个模型,采用密集和专家混合架构,参数规模从1B到8×22B不等,通过可扩展、灵活的训练流程,统一、增强并合成多样化数据集,以提升智能体在不同环境中的泛化能力与性能。实验表明,xLAM在多个智能体能力基准测试中表现卓越,尤其在伯克利函数调用排行榜上排名第一,超越GPT-4、Claude-3等众多模型的工具使用能力。通过发布xLAM系列,我们旨在推动开源大模型在自主智能体领域的性能提升,加速研究进程,并实现高性能模型在智能体任务中的普及。模型已开放获取:https://huggingface.co/collections/Salesforce/xlam-models-65f00e2a0a63bbcd1c2dade4
原文摘要 · Abstract (English)
Autonomous agents powered by large language models (LLMs) have attracted significant research interest. However, the open-source community faces many challenges in developing specialized models for agent tasks, driven by the scarcity of high-quality agent datasets and the absence of standard protocols in this area. We introduce and publicly release xLAM, a series of large action models designed for AI agent tasks. The xLAM series includes five models with both dense and mixture-of-expert architectures, ranging from 1B to 8x22B parameters, trained using a scalable, flexible pipeline that unifies, augments, and synthesizes diverse datasets to enhance AI agents' generalizability and performance across varied environments. Our experimental results demonstrate that xLAM consistently delivers exceptional performance across multiple agent ability benchmarks, notably securing the 1st position on the Berkeley Function-Calling Leaderboard, outperforming GPT-4, Claude-3, and many other models in terms of tool use. By releasing the xLAM series, we aim to advance the performance of open-source LLMs for autonomous AI agents, potentially accelerating progress and democratizing access to high-performance models for agent tasks. Models are available at https://huggingface.co/collections/Salesforce/xlam-models-65f00e2a0a63bbcd1c2dade4
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。