arXiv:2506.17209cs.CL2025-06被引 24

微调会降低大模型安全性能,且评估结果极不稳定。

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency

  • 微调使通用大模型安全对齐能力下降,即使数据无有害内容。
  • 相同实验条件下,安全评估结果波动巨大,无法可靠复现。
  • 适合关注模型安全与评估可信度的研究者阅读。

为特定领域或任务微调通用大语言模型(LLM)已成为普通用户的常规操作。然而,微调已被证实会移除模型的安全对齐特性,即便微调数据不包含任何有害内容。由于微调应用广泛且攻击本身看似无害,这一问题构成重大缺陷。多数善意开发者可能未意识到其部署的模型安全性已降低。同时,该漏洞易被恶意者利用以绕过安全防护。为解决此问题,需建立可靠、可复现的安全评估体系。本文研究安全基准在实验流程细微变化及大模型随机性下的鲁棒性,发现即使看似无关的微调设置调整,也会导致安全评估结果显著波动。该现象严重影响未来研究结果的可比性,亟需改进报告规范。

原文摘要 · Abstract (English)

Fine-tuning a general-purpose large language model (LLM) for a specific domain or task has become a routine procedure for ordinary users. However, fine-tuning is known to remove the safety alignment features of the model, even when the fine-tuning data does not contain any harmful content. We consider this to be a critical failure mode of LLMs due to the widespread uptake of fine-tuning, combined with the benign nature of the "attack". Most well-intentioned developers are likely unaware that they are deploying an LLM with reduced safety. On the other hand, this known vulnerability can be easily exploited by malicious actors intending to bypass safety guardrails. To make any meaningful progress in mitigating this issue, we first need reliable and reproducible safety evaluations. In this work, we investigate how robust a safety benchmark is to trivial variations in the experimental procedure, and the stochastic nature of LLMs. Our initial experiments expose surprising variance in the results of the safety evaluation, even when seemingly inconsequential changes are made to the fine-tuning setup. Our observations have serious implications for how researchers in this field should report results to enable meaningful comparisons in the future.

大模型安全微调风险评估可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。