arXiv:2410.15821cs.AI2024-10被引 22

微调会显著改变开源大模型的毒性输出,且效果难以预测。

The effect of fine-tuning on language model toxicity

  • 用低秩适配在非对抗数据上微调,少量参数即可改变模型毒性。
  • 开发者指令微调能降低毒性,但社区微调后毒性变化不可控。
  • 研究提醒:微调可能意外放大或抑制模型毒性,需谨慎使用。

随着开源模型普及和高效参数微调技术的发展,微调语言模型日益流行。然而,微调可能影响模型的安全性。本文评估了微调对Gemma、Llama和Phi等开源模型生成毒性内容倾向的影响。通过三项实验,对比了模型开发者在指令微调中降低毒性的效果。结果表明,在开发者已微调的模型上,仅用少量参数高效的低秩适配(LoRA)在非对抗数据集上进行微调,就能显著改变不同模型的毒性表现。最后,研究揭示了真实场景中社区贡献者微调模型后毒性率出现难以预测的偏差,凸显了微调对安全性的潜在风险。

原文摘要 · Abstract (English)

Fine-tuning language models has become increasingly popular following the proliferation of open models and improvements in cost-effective parameter efficient fine-tuning. However, fine-tuning can influence model properties such as safety. We assess how fine-tuning can impact different open models' propensity to output toxic content. We assess the impacts of fine-tuning Gemma, Llama, and Phi models on toxicity through three experiments. We compare how toxicity is reduced by model developers during instruction-tuning. We show that small amounts of parameter-efficient fine-tuning on developer-tuned models via low-rank adaptation on a non-adversarial dataset can significantly alter these results across models. Finally, we highlight the impact of this in the wild, demonstrating how toxicity rates of models fine-tuned by community contributors can deviate in hard-to-predict ways.

语言模型微调毒性控制安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。