arXiv:2412.15660cs.AIcs.CL2024-12被引 3

为企事业场景定制大模型函数调用能力,提升准确率与稳定性。

Adaptable and Precise: Enterprise-Scenario LLM Function-Calling Capability Training Pipeline

  • 基于真实业务场景生成和增强数据,微调专用函数调用模型。
  • 在人力数字化场景中构建2295份标注数据,模型精度超GPT-4。
  • 适合需要高可靠函数调用的企业级AI应用开发者使用。

企业拥有大量分散于各职能的API资产,构成现有业务流程的核心。通过将这些API作为功能工具,企业可设计多样化、场景化的智能体应用,以本地部署的函数调用模型为引擎。然而,通用模型在计算效率、输出准确性和稳定性方面难以满足企业需求,亟需针对具体场景进行适配。本文提出一套面向真实业务场景的函数调用能力训练流水线,包含场景化函数调用数据的合成与增强、模型微调及性能评估分析。在数字人力资源场景中,我们生成1,260个全自动生成样本和1,035个人工标注增强样本。采用Qwen2.5-Coder-7B-Instruct作为基底模型,基于四张24GB显存的GPU,使用LoRA方法进行微调。实验表明,该模型在测试集上的准确率超越GPT-4和GPT-4o,验证了该流水线在训练场景化函数调用模型方面的可靠性。

原文摘要 · Abstract (English)

Enterprises possess a vast array of API assets scattered across various functions, forming the backbone of existing business processes. By leveraging these APIs as functional tools, enterprises can design diverse, scenario-specific agent applications, driven by on-premise function-calling models as the core engine. However, generic models often fail to meet enterprise requirements in terms of computational efficiency, output accuracy, and stability, necessitating scenario-specific adaptation. In this paper, we propose a training pipeline for function-calling capabilities tailored to real-world business scenarios. This pipeline includes the synthesis and augmentation of scenario-specific function-calling data, model fine-tuning, and performance evaluation and analysis. Using this pipeline, we generated 1,260 fully AI-generated samples and 1,035 augmented manually-labeled samples in digital HR agent scenario. The Qwen2.5-Coder-7B-Instruct model was employed as the base model and fine-tuned using the LoRA method on four GPUs with 24GB VRAM. Our fine-tuned model demonstrated outstanding performance in evaluations and practical applications, surpassing GPT-4 and GPT-4o in accuracy on the test set. These results validate the reliability of the proposed pipeline for training scenario-specific function-calling models.

函数调用企业应用模型微调LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。