用知识驱动合成数据,让小样本也能提升大模型对话能力
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
- 两阶段生成:先构建世界知识树,再通过自我反思优化数据
- 仅20K条合成数据就超越传统方法,72B模型也有效
- 适合想低成本提升模型对齐的开发者和研究者
监督微调(SFT)数据质量对提升大语言模型(LLM)对话能力至关重要。随着模型日益先进,高质量人工标注的SFT数据成为瓶颈,亟需依赖合成数据。本文提出Condor,一种基于世界知识树与自我反思精炼的两阶段合成数据生成框架,可规模化生成高质量SFT数据。实验表明,仅用20K条Condor生成样本微调的基础模型性能即优于基线。额外的精炼阶段支持不同规模(最高达72B)模型的迭代自提升,验证了该方法的有效性。此外,对后训练中合成数据量级的探索揭示了显著未开发的性能提升潜力,为未来研究开辟新方向。
原文摘要 · Abstract (English)
The quality of Supervised Fine-Tuning (SFT) data plays a critical role in enhancing the conversational capabilities of Large Language Models (LLMs). However, as LLMs become more advanced, the availability of high-quality human-annotated SFT data has become a significant bottleneck, necessitating a greater reliance on synthetic training data. In this work, we introduce Condor, a novel two-stage synthetic data generation framework that incorporates World Knowledge Tree and Self-Reflection Refinement to produce high-quality SFT data at scale. Our experimental results demonstrate that a base model fine-tuned on only 20K Condor-generated samples achieves superior performance compared to counterparts. The additional refinement stage in Condor further enables iterative self-improvement for LLMs at various scales (up to 72B), validating the effectiveness of our approach. Furthermore, our investigation into the scaling for synthetic data in post-training reveals substantial unexplored potential for performance improvements, opening promising avenues for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。