arXiv:2606.03648cs.CLcs.AI2026-06

fine-tuning可能破坏模型安全,需以具体能力目标为基准评估

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

论文配图:Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability
图 1 · 摘自论文原文
  • 以特定能力目标为基础评估微调对安全的影响
  • 微调后模型对安全提示产生混乱输出,自动判断失效
  • 不同评测基准和评估者会得出相反结论,需谨慎选择

通过微调将基础大语言模型适配用户任务或风格时,可能损害模型安全性。以往研究在有限且看似随机的实验设置下评估微调对安全的影响。我们主张,将微调锚定于特定能力目标,是避免随意实验选择、获得有意义安全影响结论,并实现缓解方法一致性比较的关键。本文从能力与安全双维度评估微调影响,结果揭示三大问题:(1) 微调模型对安全提示可能生成不连贯输出;(2) 自动化安全判断对这类不连贯输出不可靠;(3) 关于微调影响的结论会随安全评测基准和评估者的选择而变化。

原文摘要 · Abstract (English)

Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previous works examined the effects of fine-tuning on model safety in limited and seemingly random experimental settings. We argue that anchoring fine-tuning to a specific capability goal is essential for avoiding arbitrary empirical choices, allowing us to draw meaningful conclusions about safety impacts, and to compare mitigation methods on a consistent basis. We conduct a multi-dimensional evaluation of the effects of fine-tuning on model behavior by focusing on capability as well as safety. Our results surface important issues that (1) fine-tuned models can produce incoherent generations in response to safety prompts, (2) automated safety judgments are unreliable for such incoherent outputs, and (3) the conclusions about the effects of fine-tuning can change depending on the choice of safety benchmark as well as the safety evaluator.

大模型安全微调评估能力对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。