arXiv:2601.20144cs.CL2026-01ACL被引 12

用可验证数据训练更靠谱的工具调用智能体,应对模糊、变化和不可行的用户需求。

Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents

  • 通过多轮探索生成真实交互轨迹,再转为可控任务
  • 7个主流大模型在复杂场景下普遍失败,成功率低
  • 微调轻量模型后显著提升性能,泛化能力更强

工具调用智能体正广泛应用于实际客户服务流程。然而,现有研究多聚焦于理想化、固定且明确的任务场景。现实中,用户请求常存在(1)意图模糊、(2)随时间变化或(3)因政策限制无法实现等问题,相关训练与评估数据严重不足。为此,我们提出Trajectory2Task,一个可验证的数据生成管道,用于在三种真实用户场景下大规模研究工具使用:模糊意图、意图变化和不可行意图。该管道首先通过多轮交互生成有效的工具调用轨迹,再将其转化为面向用户的任务并控制意图变化。生成的任务支持闭环评估与训练。我们在生成的复杂用户场景任务上对7个先进大模型进行评测,发现其频繁失败。利用任务回放中获得的成功轨迹,我们微调轻量级大模型,结果显示在所有三种条件下均有持续提升,并展现出对未见工具领域更强的泛化能力,表明工具调用能力显著增强。

原文摘要 · Abstract (English)

Tool-calling agents are increasingly deployed in real-world customer-facing workflows. Yet most studies on tool-calling agents focus on idealized settings with general, fixed, and well-specified tasks. In real-world applications, user requests are often (1) ambiguous, (2) changing over time, or (3) infeasible due to policy constraints, and training and evaluation data that cover these diverse, complex interaction patterns remain under-represented. To bridge the gap, we present Trajectory2Task, a verifiable data generation pipeline for studying tool use at scale under three realistic user scenarios: ambiguous intent, changing intent, and infeasible intents. The pipeline first conducts multi-turn exploration to produce valid tool-call trajectories. It then converts these trajectories into user-facing tasks with controlled intent adaptations. This process yields verifiable task that support closed-loop evaluation and training. We benchmark seven state-of-the-art LLMs on the generated complex user scenario tasks and observe frequent failures. Finally, using successful trajectories obtained from task rollouts, we fine-tune lightweight LLMs and find consistent improvements across all three conditions, along with better generalization to unseen tool-use domains, indicating stronger tool-calling ability.

工具调用大模型数据生成鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。