arXiv:2502.06589cs.CLcs.AI2025-02NAACL被引 5

用大规模数据持续预训练,提升大模型的自主执行能力。

Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training

  • 构建1030亿条专用数据,涵盖76537个API调用轨迹与文档。
  • 在三个代理基准上表现超越开源小模型,媲美商业大模型。
  • 适合研究智能体、工具调用和持续学习的开发者与学者。

由于面向智能体的预训练数据稀缺,基于大语言模型的自主智能体通常依赖复杂的提示工程或大量微调,往往无法引入新能力,同时保持强泛化性。我们提出Hephaestus-Forge,首个旨在增强大语言模型智能体基础能力的大规模预训练语料库,涵盖API函数调用、内在推理与规划、环境反馈适应等能力。该语料库包含1030亿条智能体相关数据,涉及76,537个API,既包括工具文档以引入函数知识,也包含函数调用轨迹以强化内在推理能力。为探索有效训练方案,我们研究缩放定律,确定最优的数据混合比例。通过在Hephaestus-Forge上进行持续预训练,Hephaestus在三个智能体基准测试中表现优于小型至中型开源大模型,并媲美商业大模型,验证了该预训练语料库在提升大模型智能体基础能力与泛化性方面的有效性。

原文摘要 · Abstract (English)

Due to the scarcity of agent-oriented pre-training data, LLM-based autonomous agents typically rely on complex prompting or extensive fine-tuning, which often fails to introduce new capabilities while preserving strong generalizability. We introduce Hephaestus-Forge, the first large-scale pre-training corpus designed to enhance the fundamental capabilities of LLM agents in API function calling, intrinsic reasoning and planning, and adapting to environmental feedback. Hephaestus-Forge comprises 103B agent-specific data encompassing 76,537 APIs, including both tool documentation to introduce knowledge of API functions and function calling trajectories to strengthen intrinsic reasoning. To explore effective training protocols, we investigate scaling laws to identify the optimal recipe in data mixing ratios. By continual pre-training on Hephaestus-Forge, Hephaestus outperforms small- to medium-scale open-source LLMs and rivals commercial LLMs on three agent benchmarks, demonstrating the effectiveness of our pre-training corpus in enhancing fundamental agentic capabilities and generalization of LLMs to new tasks or environments.

智能体持续预训练API调用大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。