arXiv:2511.01490cs.CL2025-11ACL

用多样合成数据微调大模型,能提升输出多样性与质量

Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning

  • 用多源合成数据微调,缓解分布坍塌问题
  • 合成数据微调后输出质量更高,但潜在风险也更大
  • 多源合成数据可有效降低自偏好偏差,效果接近真人数据

随着合成数据在语言模型开发中的广泛应用,理解其对模型行为的影响至关重要。本文研究了合成数据来源多样性对微调后大语言模型的影响,重点关注分布坍塌、对抗鲁棒性及自偏好偏差三个维度。结果表明,使用多源合成数据微调可缓解分布坍塌,保持输出分布的广度与文本多样性。尽管人类和合成数据均可消除安全防护机制,但合成数据微调后的输出质量更高,使生成内容更可用也更危险。此外,微调可减少自偏好偏差,其中真人数据最有效,其次是多源合成数据。

原文摘要 · Abstract (English)

As synthetic data becomes widely used in language model development, understanding its impact on model behavior is crucial. This paper investigates the impact of the diversity of sources of synthetic data on fine-tuned large language models. We focus on three key dimensions: distribution collapse, adversarial robustness, and self-preference bias. Our findings reveal that fine-tuning models on synthetic data from diverse sources can mitigate distribution collapse, preserving the breadth of the output distribution and the diversity of the output text. Furthermore, while both human and synthetic fine-tuning data can remove safeguards, we observe a tendency for higher output quality in the latter case, thus making outputs potentially more usable and dangerous. Finally, we also find evidence that fine-tuning reduces self-preference bias, with human data being the most effective, followed by multi-source synthetic data.

合成数据大模型微调多样性输出质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。