arXiv:2605.04572cs.AIcs.LG2026-05中稿 · ICML

通过分析参数动态变化,量化微调中每个样本的安全风险。

From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning

论文配图:From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning
图 1 · 摘自论文原文
  • 基于参数更新方向差异,计算样本对安全性的潜在危害。
  • 实验证明该方法能有效识别高风险训练样本。
  • 适用于不同模型架构与微调方法,具有强泛化能力。

大型语言模型(LLM)的安全对齐极为脆弱,仅用少量良性样本微调就可能抹去模型从数百万条偏好数据中学到的安全行为。现有研究通过对比微调前后参数和隐藏状态解释此现象,但忽略了参数在微调过程中的动态演变。本文通过分析参数动态,揭示了安全退化的关键机制:良性微调会导致参数持续向危险方向漂移,逐步削弱模型安全性。这一发现表明,导致更大漂移的样本具有更高的微调风险。基于此,我们提出样本级安全退化量化方法(SQSD),通过测量每个样本引起的参数更新在危险与安全方向上的投影差异,连续生成风险评分。在多个模型和数据集上的实验表明,SQSD能有效量化样本级微调风险,并在不同模型架构、参数规模及参数高效方法间展现出良好可迁移性。

原文摘要 · Abstract (English)

Safety alignment of Large Language Models (LLMs) is extremely fragile, as fine-tuning on a small number of benign samples can erase safety behaviors learned from millions of preference examples. Existing studies attempt to explain this phenomenon by comparing parameters and hidden states before and after fine-tuning, but overlook their dynamic evolution during fine-tuning. In this paper, we uncover a critical mechanism underlying safety degradation by analyzing parameter dynamics, where benign fine-tuning causes parameters to cumulatively drift toward danger-aligned directions, progressively undermining the model's safety. This finding suggests that samples contributing more to this drift has greater fine-tuning risks. Based on this insight, we propose a method of Sample-Level Quantification of Safety Degradation (SQSD), which quantifies the influence of each training sample on safety degradation. Specifically, SQSD computes continuous risk scores to samples by measuring their induced parameter updates' projection difference between danger and safety directions. Extensive experiments across multiple models and datasets demonstrate that SQSD effectively quantifies sample-level fine-tuning risks and exhibits strong transferability across model architectures, parameter scales, and parameter-efficient methods.

安全对齐微调风险参数动态风险评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。