arXiv:2501.09766cs.CLcs.AI2025-01EMNLP被引 6

通过动态缺陷校准提升大模型工具使用能力

iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use

  • 用蒙特卡洛树搜索增强合成数据多样性
  • 迭代定位并优化模型缺陷,性能提升6.5%
  • 适合需要复杂工具调用的智能助手研发

将外部工具赋予大语言模型是提升其能力的有前景方法,尤其在复杂任务中。通过真实世界模拟生成工具使用数据是一种有效途径。然而研究发现,随着合成数据量增加,训练收益显著下降,模型难以从更多数据中获益,无法在复杂场景中获得高级工具使用能力。我们发现该限制通常表现为响应中的片段缺失(即参数错误)。为此,提出一种迭代强化微调策略:(1) 通过蒙特卡洛树搜索的路径探索增强合成数据响应的多样性;(2) 通过构建细粒度偏好对迭代定位模型缺陷,并利用偏好优化算法实现针对性改进。实验表明,该方法相比同规模基线模型性能提升13.11%,在复杂场景下较基线提升6.5%,且优于更大规模的开源与闭源模型。

原文摘要 · Abstract (English)

Augmenting large language models (LLMs) with external tools is a promising approach to enhance their capabilities, especially for complex tasks. Synthesizing tool-use data through real-world simulations is an effective way to achieve this. However, our investigation reveals that training gains significantly decay as synthetic data increases. The model struggles to benefit from additional synthetic data, which fails to endow it with advanced tool-use capabilities in complex scenarios Moreover, we discovered that the above limitation usually manifests as a fragment deficiency (i.e., parameter errors) in response. To this end, we propose an iterative reinforced fine-tuning strategy designed to alleviate this limitation. This strategy involves: (1) enhancing the diversity of response for synthetic data through path exploration of Monte Carlo Tree Search. (2) iteratively pinpointing the model's deficiency by constructing fine-grained preference pairs, and then improving it by preference optimization algorithms for targeted improvement. The experiments show that our method achieves 13.11% better performance than the same-size base model. It achieves an improvement of 6.5% in complex scenarios compared to the baseline, and it also outperforms larger open-source and closed-source models.

工具使用强化微调大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。