arXiv:2505.24165cs.CL2025-05ACL

用标签注入实现高效指令演化,提升数据多样性与挑战性。

Tag-Evol: Achieving Efficient Instruction Evolving via Tag Injection

  • 通过注入不同组合的标签,实现可控指令演化。
  • 在多个基准上生成的数据质量显著优于现有方法。
  • 适合需要高质量多样化训练数据的研究者使用。

Evol-Instruct 在多个领域作为数据合成方法取得了显著进展。现有方法通常依赖固定策略进行演化,需人工设计且形式单一;同时迭代演化导致获取困难样本成本高昂。为此,我们提出 Tag-Evol 框架,一种更高效、多样化的指令演化方法。具体而言,Tag-Evol 利用多样且特定的知识标签作为策略,通过将不同标签组合注入原始指令,实现受控演化。在多种骨干模型和跨领域基准上的实验表明,该方法生成的演化数据显著优于其他方法。此外,我们对演化数据进行了深入分析,证实 Tag-Evol 不仅高效,还能生成更多样化、更具挑战性的数据。

原文摘要 · Abstract (English)

Evol-Instruct has made significant improvements as a data synthesis method in several areas. Existing methods typically rely on a fixed set of strategies to evolve, which require manual design and are monolithic in form. In addition, iterative evolution also makes the acquisition of hard samples expensive. In view of this, we propose the Tag-Evol framework, a more diverse and efficient instruction evolving method. Specifically, Tag-Evol uses diverse and specific knowledge tags as strategies to achieve controlled evolution by injecting different combinations of tags into the original instructions. Experiments with multiple backbones in diverse domain benchmarks show that the proposed method generates significantly better evolved data than other methods. Furthermore, we conduct a thorough analysis of the evolved data, demonstrating that Tag-Evol is not only efficient but also generates more diverse and challenging data.

指令演化数据合成标签注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。