用约束验证法生成高质量交互工具使用数据,提升模型性能。
CoVe: Training Interactive Tool-Use Agents via Constraint-Guided Verification
- 通过显式任务约束生成复杂且正确的操作轨迹。
- 43.0%和59.4%的成功率超越同规模基线,媲美17倍大的模型。
- 适合研究交互式代理、具身智能与自动化工具使用的学者。
构建多轮交互式工具使用代理面临挑战,因真实用户需求常复杂且模糊,而代理需执行确定性动作以满足需求。为此,我们提出CoVe(Constraint-Verification)——一种后训练数据合成框架,用于训练交互式工具使用代理,同时确保数据复杂性与正确性。CoVe首先定义显式任务约束,该约束兼具双重作用:引导生成复杂轨迹,并作为确定性验证器评估轨迹质量。这使得高质训练轨迹可用于监督微调(SFT),并为强化学习(RL)生成精确奖励信号。在挑战性τ²-bench基准上的评估表明,紧凑的CoVe-4B模型在航空与零售领域分别达到43.0%和59.4%的成功率;其整体性能显著优于同规模强基线,且媲美最大达其17倍大小的模型。结果表明,CoVe为生成顶尖交互式工具使用代理的训练数据提供了一条高效路径。为支持后续研究,我们开源代码、训练模型及全部12K条高质量训练轨迹。
原文摘要 · Abstract (English)
Developing multi-turn interactive tool-use agents is challenging because real-world user needs are often complex and ambiguous, yet agents must execute deterministic actions to satisfy them. To address this gap, we introduce \textbf{CoVe} (\textbf{Co}nstraint-\textbf{Ve}rification), a post-training data synthesis framework designed for training interactive tool-use agents while ensuring both data complexity and correctness. CoVe begins by defining explicit task constraints, which serve a dual role: they guide the generation of complex trajectories and act as deterministic verifiers for assessing trajectory quality. This enables the creation of high-quality training trajectories for supervised fine-tuning (SFT) and the derivation of accurate reward signals for reinforcement learning (RL). Our evaluation on the challenging $τ^2$-bench benchmark demonstrates the effectiveness of the framework. Notably, our compact \textbf{CoVe-4B} model achieves success rates of 43.0\% and 59.4\% in the Airline and Retail domains, respectively; its overall performance significantly outperforms strong baselines of similar scale and remains competitive with models up to $17\times$ its size. These results indicate that CoVe provides an effective and efficient pathway for synthesizing training data for state-of-the-art interactive tool-use agents. To support future research, we open-source our code, trained model, and the full set of 12K high-quality trajectories used for training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。