自动生成可扩展的智能体任务,解决数据稀缺难题
TaskCraft: Automated Generation of Agentic Tasks
- 通过深度与广度扩展生成多工具、分层复杂任务
- 构建约3.6万条任务数据集,支持智能体训练与评估
- 适合研究智能体训练、提示优化与自动评测的学者
智能体任务要求多步求解、自主决策、工具使用和动态推理,正成为NLP与AI发展的核心。然而现有指令数据缺乏工具交互,现有智能体基准依赖高成本人工标注,限制了可扩展性。我们提出 extsc{TaskCraft},一种自动化工作流,用于生成可调节难度、多工具、可验证执行轨迹的智能体任务。通过基于深度和宽度的扩展方式,生成结构与层次复杂的任务。实验表明,这些任务能提升生成流程中的提示优化效果,并增强对智能体基础模型的监督微调。我们构建了一个大规模合成数据集,包含约36,000个不同难度的任务,以支持未来智能体调优与评估研究。
原文摘要 · Abstract (English)
Agentic tasks, which require multi-step problem solving with autonomy, tool use, and adaptive reasoning, are becoming increasingly central to the advancement of NLP and AI. However, existing instruction data lacks tool interaction, and current agentic benchmarks rely on costly human annotation, limiting their scalability. We introduce \textsc{TaskCraft}, an automated workflow for generating difficulty-scalable, multi-tool, and verifiable agentic tasks with execution trajectories. TaskCraft expands atomic tasks using depth-based and width-based extensions to create structurally and hierarchically complex challenges. Empirical results show that these tasks improve prompt optimization in the generation workflow and enhance supervised fine-tuning of agentic foundation models. We present a large-scale synthetic dataset of approximately 36,000 tasks with varying difficulty to support future research on agent tuning and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。