用动态噪声优化让大模型自我生成数据持续进化
Dynamic Noise Preference Optimization: Self-Improvement of Large Language Models with Self-Synthetic Data
- 通过动态标注与可训练噪声注入构建偏好对
- 在多个基准上持续提升性能,生成数据质量提升29.4%
- 适合想减少人工标注、实现模型自进化的研究者
尽管大语言模型已取得显著进展,但其对大量人工标注数据的依赖限制了进一步扩展。在此背景下,利用自生成合成数据进行微调成为避免大规模人工标注的关键。然而,现有方法常无法保证迭代中的持续改进,性能在少量更新后即停滞。为此,我们提出动态噪声偏好优化(DNPO),结合动态样本标注构建偏好对,并在偏好优化中引入可控可训练的噪声注入。该方法有效防止性能停滞,实现持续提升。在Llama-3.2-3B和Zephyr-7B上的实验表明,DNPO在多个基准上均优于现有方法。此外,使用Zephyr-7B时,模型生成数据质量显著提升,在GPT-4评估中表现出29.4%的胜率差距。
原文摘要 · Abstract (English)
Although LLMs have achieved significant success, their reliance on large volumes of human-annotated data has limited their potential for further scaling. In this situation, utilizing self-generated synthetic data has become crucial for fine-tuning LLMs without extensive human annotation. However, current methods often fail to ensure consistent improvements across iterations, with performance stagnating after only minimal updates. To overcome these challenges, we introduce Dynamic Noise Preference Optimization (DNPO), which combines dynamic sample labeling for constructing preference pairs with controlled, trainable noise injection during preference optimization. Our approach effectively prevents stagnation and enables continuous improvement. In experiments with Llama-3.2-3B and Zephyr-7B, DNPO consistently outperforms existing methods across multiple benchmarks. Additionally, with Zephyr-7B, DNPO shows a significant improvement in model-generated data quality, with a 29.4% win-loss rate gap compared to the baseline in GPT-4 evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。