arXiv:2502.04511cs.CL2025-02Conference of the …被引 1

用参考样本指导数据生成,突破合成数据质量瓶颈。

Beyond Sample-Level Feedback: Using Reference-Level Feedback to Guide Data Synthesis

  • 从精选参考样本中提取优质特征,引导生成更优指令回复对
  • 生成1万条数据,使小模型在AlpacaEval上达43.96%胜率
  • 方法通用、高效,适合想提升微调数据质量的研究者

高质量指令微调数据对训练能完成真实任务的大型语言模型至关重要。尽管合成数据生成可规模化构建数据集,但其质量受限于生成模型本身。为此,本文提出参考级反馈(Reference-Level Feedback)机制,从精心筛选的参考样本中提取理想特征,用于指导生成更高品质的指令-响应对。基于此方法,我们构建了包含10,000条指令-响应对的REFED数据集。在该数据集上微调Llama-3.1-8B-Instruct和Mistral-7B-Instruct模型,均达到同规模模型中的最先进性能,尤其在AlpacaEval 2.0的长度控制测试中实现43.96%的胜率。大量实验表明,该方法持续优于传统样本级反馈,具备跨模型架构泛化能力,且以低成本生成高质量、多样化数据。

原文摘要 · Abstract (English)

High-quality instruction-tuning data is crucial for developing Large Language Models (LLMs) that can effectively navigate real-world tasks and follow human instructions. While synthetic data generation offers a scalable approach for creating such datasets, it imposes a quality ceiling where models trained on the data cannot outperform the LLM generating it. To overcome this limitation, we introduce Reference-Level Feedback, a paradigm that extracts desirable characteristics from carefully curated reference samples to guide the synthesis of higher-quality instruction-response pairs. Using this approach, we synthesize REFED, a dataset of 10K instruction-response pairs. Fine-tuning Llama-3.1-8B-Instruct and Mistral-7B-Instruct on REFED demonstrate state-of-the-art performance among similarly sized models, notably reaching a 43.96\% length-controlled win-rate on AlpacaEval 2.0. Extensive experiments demonstrate that Reference-Level Feedback consistently outperforms traditional sample-level feedback methods, generalizes across model architectures, and produces high-quality and diverse data at low cost.

数据合成指令微调模型优化参考反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。