让大模型更精准地使用工具,通过细粒度优化调用过程
TTPA: Token-level Tool-use Preference Alignment Training Framework with Fine-grained Evaluation
- 反向生成构建高质量对话数据,提升工具使用上下文
- 在三个数据集上显著提升工具调用性能,泛化能力强
- 基于错误量化评分,适合需要高精度工具使用的场景
现有工具学习方法多依赖监督微调,常忽视内部工具调用细节的精细化优化,导致偏好对齐与错误识别能力受限。为此,我们提出令牌级工具使用偏好对齐训练框架(TTPA),一种用于构建细粒度工具使用偏好数据集的训练范式,通过新颖的错误导向评分机制使大语言模型与细粒度偏好对齐。TTPA 首次引入反向数据集构建法,通过逆向生成流程创建高质量、多轮次工具使用数据;提出令牌级偏好采样(TPS),通过建模生成过程中的令牌级差异捕捉细粒度偏好;为缓解评分偏差,设计错误导向评分机制(ESM),量化工具调用错误并作为训练信号。在三个多样化基准数据集上的大量实验表明,TTPA 显著提升工具使用性能,并展现出跨模型与跨数据集的强大泛化能力。
原文摘要 · Abstract (English)
Existing tool-learning methods usually rely on supervised fine-tuning, they often overlook fine-grained optimization of internal tool call details, leading to limitations in preference alignment and error discrimination. To overcome these challenges, we propose Token-level Tool-use Preference Alignment Training Framework (TTPA), a training paradigm for constructing token-level tool-use preference datasets that align LLMs with fine-grained preferences using a novel error-oriented scoring mechanism. TTPA first introduces reversed dataset construction, a method for creating high-quality, multi-turn tool-use datasets by reversing the generation flow. Additionally, we propose Token-level Preference Sampling (TPS) to capture fine-grained preferences by modeling token-level differences during generation. To address biases in scoring, we introduce the Error-oriented Scoring Mechanism (ESM), which quantifies tool-call errors and can be used as a training signal. Extensive experiments on three diverse benchmark datasets demonstrate that TTPA significantly improves tool-using performance while showing strong generalization ability across models and datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。