arXiv:2510.21885cs.CLcs.AI2025-10被引 1

通过智能选样防止大模型微调时遗忘安全行为

Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning

  • 按指令响应行为和有害类别多样性筛选安全样本
  • 仅用0.5%额外数据,有害输出减少41%
  • 适合需要高效提升模型安全性的实践者

大型语言模型在微调过程中常因学习良性数据而遗忘原有的安全行为,这种现象称为灾难性遗忘。已有研究显示加入随机安全样本可缓解该问题,但尚不清楚哪些样本最有效。本文提出一种行为感知采样框架,根据指令-响应行为(如拒绝或遵从)以及不同有害类别间的语义多样性来选择安全样本。系统评估表明,该方法显著降低有害输出,同时保持模型有用性,仅使用0.5%的额外训练数据即实现高达41%的有害性下降。结果说明,针对性的数据选择能大幅提升大规模微调的安全性和效率。

原文摘要 · Abstract (English)

Large language models often lose previously aligned safety behaviors when fine-tuned on benign data, a phenomenon known as catastrophic forgetting. Prior work shows that adding random safety examples can mitigate this effect, but it remains unclear which examples are most effective. We propose a behavior-aware sampling framework that selects safety examples based on two complementary factors: instruction-response behavior (e.g., refusal versus compliance) and semantic diversity across harm categories. Systematic evaluation shows that this approach substantially reduces harmful outputs while maintaining helpfulness, achieving up to a 41% reduction in harmfulness with only 0.5% additional training data. These results highlight how targeted data selection can improve the safety and efficiency of fine-tuning at scale.

安全微调灾难性遗忘数据采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。