arXiv:2506.03850cs.LG2025-06ICML被引 18

识别安全数据遗忘弱点,分组对抗训练提升模型抗有害微调能力

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning

  • 基于数据遗忘脆弱性分组,动态选择低表现组样本强化学习
  • 在4个微调任务中降低有害得分,同时保持下游性能
  • 适合关注大模型安全对齐的从业者与研究者

有害微调(HFT)直接作用于开源大模型或通过微调即服务方式实施,会破坏安全对齐并带来重大威胁。现有方法通常在对齐数据上学习鲁棒表征或使有害数据不可学习,但均将每个数据样本视为同等重要,忽视了数据脆弱性模式。本文揭示,在不同微调任务中,对齐数据的某些子集始终更易发生遗忘。受此启发,我们提出脆弱性感知对齐(VAA),通过估计数据脆弱性,将数据划分为“脆弱”与“不易脆弱”两组,并采用群体分布鲁棒优化(Group DRO)框架促进平衡学习。具体而言,VAA学习一个对抗采样器,在训练中从当前表现较差的组中采样,并施加组依赖的对抗扰动,以推动各组间学习均衡。在四个微调任务上的实验表明,VAA显著降低有害得分,同时保持下游任务性能,优于当前最优基线。

原文摘要 · Abstract (English)

Harmful fine-tuning (HFT), performed directly on open-source LLMs or through Fine-tuning-as-a-Service, breaks safety alignment and poses significant threats. Existing methods aim to mitigate HFT risks by learning robust representation on alignment data or making harmful data unlearnable, but they treat each data sample equally, leaving data vulnerability patterns understudied. In this work, we reveal that certain subsets of alignment data are consistently more prone to forgetting during HFT across different fine-tuning tasks. Inspired by these findings, we propose Vulnerability-Aware Alignment (VAA), which estimates data vulnerability, partitions data into "vulnerable" and "invulnerable" groups, and encourages balanced learning using a group distributionally robust optimization (Group DRO) framework. Specifically, VAA learns an adversarial sampler that samples examples from the currently underperforming group and then applies group-dependent adversarial perturbations to the data during training, aiming to encourage a balanced learning process across groups. Experiments across four fine-tuning tasks demonstrate that VAA significantly reduces harmful scores while preserving downstream task performance, outperforming state-of-the-art baselines.

安全对齐微调防御大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。