arXiv:2604.27019cs.LGcs.CL2026-04被引 2

动态对抗微调重塑模型拒绝行为的几何结构,平衡安全与可用性。

Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry

  • 通过动态对抗训练重新组织拒绝控制的低维几何载体。
  • 早期检查点拒绝攻击成功率归零,但误拒严重;后期部分恢复可用性,攻击成功率升至61.3%。
  • 揭示了安全与实用性的权衡机制,适合关注对齐鲁棒性的研究者。

安全对齐的语言模型需拒绝有害请求而不过度拒绝,但动态对抗微调如何改变拒绝控制的载体仍不明确:是KL约束方向,还是不引起显著安全提示分布偏移的小子空间?我们研究了一个70亿参数模型,在监督微调(SFT)和鲁棒拒绝动态防御(R2D2)下,结合HarmBench、StrongREJECT、XSTest评估与五锚点几何测量、因果干预及稀疏自适应压力测试。R2D2在早期检查点使固定源HarmBench攻击成功率降至0,但此时XSTest拒绝率最高且通过不了良性可用性审计。后期检查点部分恢复实用性,同时攻击成功率回升,自适应GCG攻击成功率在第250步达0.415,第500步达0.613。内部分析显示,R2D2在前100步保持深层可接受的拒绝控制载体,之后将其转移至浅层;而SFT更早迁移,但鲁棒性较弱。有效秩维持在1.24左右,SFT主角度漂移更大,表明维度扩张与漂移幅度不足以解释现象。因果干预支持一个低维但与实用性耦合的载体。结果支持R2D2在鲁棒性-实用性边界上的几何重构观点,但未建立自适应鲁棒性。

原文摘要 · Abstract (English)

Safety-aligned language models must refuse harmful requests without broad over-refusal, but it remains unclear how dynamic adversarial fine-tuning changes refusal-control carriers: Kullback--Leibler (KL)-constrained directions or small subspaces that causally modulate refusal without large safe-prompt distribution shifts. We study a 7B backbone under supervised fine-tuning (SFT) and Robust Refusal Dynamic Defense (R2D2), aligning HarmBench, StrongREJECT, and XSTest evaluations with five-anchor geometry measurements, causal interventions, and sparse adaptive stress tests. R2D2 drives fixed-source HarmBench attack success to zero at early checkpoints; however, these checkpoints also exhibit maximal XSTest refusal and fail a benign-utility audit. Later checkpoints partially recover utility-facing behavior while reopening attack success, with adaptive GCG attack success rate rising to 0.415 at step 250 and 0.613 at step 500. Internally, R2D2 preserves a late-layer admissible refusal-control carrier through step 100 and then relocates the best admissible carrier to an early layer; SFT relocates earlier yet remains less robust. Effective rank stays near 1.24, and SFT shows larger principal-angle drift, arguing against both dimensional expansion and drift magnitude as sufficient explanations. Causal interventions support a low-dimensional but utility-coupled carrier. These results support a geometry-reorganization account of R2D2 along a robustness--utility frontier, without establishing adaptive robustness.

模型对齐对抗训练拒绝机制几何分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。