用技能图谱自动生成多样化终端任务,提升智能体训练效果
Toward Scalable Terminal Task Synthesis via Skill Graphs

- 基于场景中介的技能图谱生成任务路径
- 在Terminal-Bench上验证有效,支持多任务多样性
- 适合需要高质量终端指令训练的智能体研究者
终端智能体在自主命令行执行方面展现出巨大潜力,但其训练受限于高质量、多样化执行轨迹的稀缺。现有方法通过合成大规模终端任务实例来缓解这一瓶颈,但主要关注任务数量扩展,对训练中实际体验的轨迹多样性控制有限。本文提出SkillSynth,一种基于场景中介技能图谱的自动化终端任务合成框架。该框架首先构建大规模技能图谱,其中场景作为中间过渡节点连接多样命令行技能;随后从图中采样路径作为真实工作流的抽象,并利用多智能体环境将其实例化为可执行任务。通过图采样路径实现对最小执行轨迹多样性的显式控制。在Terminal-Bench上的实验验证了SkillSynth的有效性。此外,由SkillSynth合成的任务实例已用于训练Hy3 Preview,显著提升了其在终端场景下的智能体能力。
原文摘要 · Abstract (English)
Terminal agents have demonstrated strong potential for autonomous command-line execution, yet their training remains constrained by the scarcity of high-quality and diverse execution trajectories. Existing approaches mitigate this bottleneck by synthesizing large-scale terminal task instances for trajectory sampling. However, they primarily focus on scaling the number of tasks while providing limited control over the diversity of execution trajectories that agents actually experience during training. In this paper, we present SkillSynth, an automated framework for terminal task synthesis built on a scenario-mediated skill graph. SkillSynth first constructs a large-scale skill graph, where scenarios serve as intermediate transition nodes that connect diverse command-line skills. It then samples paths from this graph as abstractions of real-world workflows, and uses a multi-agent harness to instantiate them into executable task instances. By grounding task synthesis in graph-sampled workflow paths, SkillSynth explicitly controls the diversity of minimal execution trajectories required to solve the synthesized tasks. Experiments on Terminal-Bench demonstrate the effectiveness of SkillSynth. Moreover, task instances synthesized by SkillSynth have been adopted to train Hy3 Preview, contributing to its enhanced agentic capabilities in terminal-based settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。