10%错误数据就让大模型性能暴跌,50%以上正确率才可能恢复
How Much of Your Data Can Suck? Thresholds for Domain Performance and Emergent Misalignment in LLMs
- 用不同比例错误数据微调gpt-4o,测试其在编码、金融等领域的表现
- 10%-25%错误数据即导致性能大幅下降,需超50%正确数据才能恢复
- 高风险场景应慎用微调,直接用基础模型更安全可靠
本文研究错误数据对大语言模型(如gpt-4o)在监督微调(SFT)中性能与安全的影响。尽管LLM广泛应用于金融、编程、医疗、法律等领域,但使用错误数据可能导致‘涌现性偏差’,产生有害或欺骗性输出。我们评估了在编码、金融、健康、法律四个领域中,以10%至90%比例混入明显或隐晦错误数据的gpt-4o模型表现。结果表明,仅10%-25%错误数据便显著降低任务性能,而道德对齐未受影响;至少需50%正确数据才能稳定恢复性能,且极少能超越基座模型的鲁棒性与安全性——后者无需微调即实现近乎完美的对齐,且无危险输出。研究强调错误数据代价高昂,提示高价值应用应优先保证极高质量的数据筛选,或直接使用强健基座模型避免不必要的微调。
原文摘要 · Abstract (English)
This paper investigates the impact of incorrect data on the performance and safety of large language models (LLMs), specifically gpt-4o, during supervised fine-tuning (SFT). Although LLMs become increasingly vital across broad domains like finance, coding, law, and health, fine-tuning on incorrect data can lead to "emergent misalignment," producing harmful or deceptive outputs unrelated to the intended task. We evaluate gpt-4o models fine-tuned with varying ratios (10\% to 90\% correct) of both obviously and subtly incorrect data across four domains: coding, finance, health, and legal. Our findings show that even modest amounts of incorrect data (10-25\%) dramatically degrade domain performance and not moral alignment. A clear threshold of at least 50\% correct data is needed for models to consistently recover strong performance, though they rarely match the robustness and safety of the base model, which exhibits near-perfect alignment and zero dangerous completions out-of-the-box. This research emphasizes that the cost of incorrect data is heavy, highlighting the critical need for extremely high-quality data curation or, alternatively, leveraging robust base models without unnecessary fine-tuning for high-stakes applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。