无需人工标注,自动生成领域专用指令数据提升大模型专业能力
DS$^2$-Instruct: Domain-Specific Data Synthesis for Large Language Models Instruction Tuning
- 用任务关键词+布卢姆认知层级生成多样化指令
- 在数学、金融等7个领域上效果优于现有方法
- 零样本生成+自洽验证,适合垂直领域模型训练
将大语言模型适配到特定领域需要高质量的指令微调数据,但人工标注成本高昂。现有数据合成方法多针对通用任务,难以捕捉领域术语和推理模式。为此,我们提出DS²-Instruct,一个无需人工监督的零样本框架,用于生成领域专用指令数据。该方法首先生成任务相关关键词以确保领域覆盖全面;然后通过将关键词与布卢姆认知层级中的不同认知水平组合,生成多样化的指令;最后利用自洽性验证保证数据质量。我们在数学、金融、逻辑推理等七个挑战性领域应用该框架生成数据集。综合评估表明,基于我们生成数据微调的模型,在性能上显著优于现有数据生成方法。
原文摘要 · Abstract (English)
Adapting Large Language Models (LLMs) to specialized domains requires high-quality instruction tuning datasets, which are expensive to create through human annotation. Existing data synthesis methods focus on general-purpose tasks and fail to capture domain-specific terminology and reasoning patterns. To address this, we introduce DS$^2$-Instruct, a zero-shot framework that generates domain-specific instruction datasets without human supervision. Our approach first generates task-informed keywords to ensure comprehensive domain coverage. It then creates diverse instructions by pairing these keywords with different cognitive levels from Bloom's Taxonomy. Finally, it uses self-consistency validation to ensure data quality. We apply this framework to generate datasets across seven challenging domains, such as mathematics, finance, and logical reasoning. Comprehensive evaluation demonstrates that models fine-tuned on our generated data achieve substantial improvements over existing data generation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。