用大模型自动生成高质量指令数据,无需外部资源。
DecIF: Improving Instruction-Following through Meta-Decomposition
- 通过元分解迭代生成多样化指令信息,提升结构与语义质量。
- 将指令拆解为原子级评估标准,有效剔除错误配对。
- 全自主框架,适合需要海量指令数据的训练场景。
指令遵循已成为大语言模型的关键能力。现有方法通常依赖预存文档或外部资源合成指令遵循数据,限制了灵活性与泛化能力。本文提出DecIF,一种完全自主、基于元分解引导的框架,仅使用大语言模型即可生成多样且高质量的指令遵循数据。在指令生成阶段,引导大模型迭代生成多种元信息,并结合响应约束形成结构良好、语义丰富的指令;进一步利用大模型检测并修复生成指令中的潜在不一致。在响应生成阶段,将每条指令分解为原子级评估标准,实现严格验证并剔除不准确的指令-响应对。广泛实验表明,DecIF在多种场景与设置下均表现优异。进一步分析显示其具备强灵活性、可扩展性与泛化能力,能自动合成高质量指令数据。
原文摘要 · Abstract (English)
Instruction-following has emerged as a crucial capability for large language models (LLMs). However, existing approaches often rely on pre-existing documents or external resources to synthesize instruction-following data, which limits their flexibility and generalizability. In this paper, we introduce DecIF, a fully autonomous, meta-decomposition guided framework that generates diverse and high-quality instruction-following data using only LLMs. DecIF is grounded in the principle of decomposition. For instruction generation, we guide LLMs to iteratively produce various types of meta-information, which are then combined with response constraints to form well-structured and semantically rich instructions. We further utilize LLMs to detect and resolve potential inconsistencies within the generated instructions. Regarding response generation, we decompose each instruction into atomic-level evaluation criteria, enabling rigorous validation and the elimination of inaccurate instruction-response pairs. Extensive experiments across a wide range of scenarios and settings demonstrate DecIF's superior performance on instruction-following tasks. Further analysis highlights its strong flexibility, scalability, and generalizability in automatically synthesizing high-quality instruction data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。