用强化学习让小模型也能精准调用工具
Advancing SLM Tool-Use Capability using Reinforcement Learning
- 用分组相对策略优化提升小模型工具调用能力
- 在真实场景中显著提高函数调用和参数使用准确率
- 适合资源受限环境下部署智能代理的开发者
在工具增强型AI代理日益重要的背景下,本研究发现,分组相对策略优化(GRPO)能够有效提升小型语言模型(SLMs)的工具使用能力。传统上SLMs在工具调用方面存在局限,而大模型(LLMs)已具备通过外部数据与内部资源扩展功能的能力。随着代理系统日趋复杂,工具使用成为关键能力。尽管LLMs已有显著进展,但SLMs在资源受限场景下仍难以准确整合工具使用。本文通过设计包含结构化JSON输出、正确工具选择和精确参数使用的奖励机制,验证了GRPO能显著提升SLMs的工具使用准确性(函数调用/JSON输出)。该方法提供了一种计算高效的训练方案,有利于提升SLMs在真实应用中的可部署性。
原文摘要 · Abstract (English)
In an era where tool-augmented AI agents are becoming increasingly vital, our findings highlight the ability of Group Relative Policy Optimization (GRPO) to empower SLMs, which are traditionally constrained in tool use. The ability to use tools effectively has become a defining feature of Large Language Models (LLMs), allowing them to access external data and internal resources. As AI agents grow more sophisticated, tool-use capabilities have become indispensable. While LLMs have made significant progress in this area, Small Language Models (SLMs) still face challenges in accurately integrating tool use, especially in resource-constrained settings. This study investigates how Reinforcement Learning, specifically Group Relative Policy Optimization (GRPO), can enhance the tool-use accuracy of SLMs. By designing a well-defined reward system that reinforces structured JSON output, correct tool selection, and precise parameter usage, we demonstrate that GRPO enables SLMs to achieve significant improvements in tool-use capabilities (function calling/JSON output). Our approach provides a computationally efficient training method that enhances SLMs practical deployment in real-world AI applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。