arXiv:2412.16469cs.CL2024-12被引 4

顺序训练导致模型更忘安全知识,且对特定群体影响更大

Chained Tuning Leads to Biased Forgetting

  • 按下游任务顺序微调会加剧安全知识遗忘
  • 安全信息遗忘程度高于常规微调,且对部分群体更严重
  • 提出新指标‘偏见性遗忘’,适合持续学习场景研究者

大型语言模型在下游任务上微调时,常会削弱先前训练中习得的能力,这种现象称为灾难性遗忘,对部署模型的安全性有重要影响。本文首先发现:按下游任务顺序微调的模型,比反向顺序微调的模型遗忘其安全调控更严重;其次,遗忘对特定群体的安全信息影响尤为显著。为量化该现象,我们提出新的度量标准——偏见性遗忘。通过系统评估任务顺序对遗忘的影响,并应用缓解策略帮助模型恢复遗忘内容。研究成果有助于指导大模型在持续学习场景下的链式微调方法,实现更安全、更低毒性的模型训练。

原文摘要 · Abstract (English)

Large language models (LLMs) are often fine-tuned for use on downstream tasks, though this can degrade capabilities learned during previous training. This phenomenon, often referred to as catastrophic forgetting, has important potential implications for the safety of deployed models. In this work, we first show that models trained on downstream tasks forget their safety tuning to a greater extent than models trained in the opposite order. Second, we show that forgetting disproportionately impacts safety information about certain groups. To quantify this phenomenon, we define a new metric we term biased forgetting. We conduct a systematic evaluation of the effects of task ordering on forgetting and apply mitigations that can help the model recover from the forgetting observed. We hope our findings can better inform methods for chaining the finetuning of LLMs in continual learning settings to enable training of safer and less toxic models.

大模型微调灾难性遗忘安全性偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。