arXiv:2605.01913cs.LGcs.AI2026-05被引 1

提出几何保持微调方法,让大模型在任务适配中仍能拒绝有害请求。

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

论文配图:RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
图 1 · 摘自论文原文
  • 在隐藏表示空间约束更新,保护安全特征的几何结构。
  • 在多个模型上测试,攻击成功率接近原始安全模型。
  • 适合需要安全与性能兼顾的应用场景。

为下游任务微调安全对齐的语言模型常导致拒绝行为严重退化,使模型易受恶意滥用。尽管先前研究发现安全相关特征编码于模型激活空间的结构化表征中,但这些表征在微调过程中的变化机制及对齐退化的根源仍不明确。本文分析表明,标准微调会引发安全相关表征的系统性漂移,扭曲其几何结构,并引入任务优化与安全特征间的干扰,共同导致有害顺从性上升。基于此,我们提出REFUSALGUARD——一种在表示层面保持安全结构的微调框架。该方法通过约束隐藏表示空间的更新,确保安全中介组件稳定,同时允许在互补方向上进行任务特异性学习。我们在包括LLaMA、Gemma、Qwen在内的多个模型家族上评估了REFUSALGUARD,在AdvBench、DirectHarm4、JailbreakBench等对抗性安全基准以及下游任务上表现优异,攻击成功率与基础安全对齐模型相当,且任务性能显著优于基线。

原文摘要 · Abstract (English)

Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse. While prior work has shown that safety-relevant features are encoded in structured representations within the model's activation space, how these representations change during fine-tuning and why alignment degrades remains poorly understood. In this work, we investigate the representation-level mechanisms underlying alignment degradation. Our analysis shows that standard fine-tuning induces systematic drift in safety-relevant representations, distorts their geometric structure, and introduces interference between task optimization and safety features. These effects collectively lead to increased harmful compliance. Motivated by these findings, we introduce REFUSALGUARD, a representation-level fine-tuning framework that preserves safety-relevant structure during model adaptation. Our approach constrains updates in hidden representation space, ensuring that safety-mediating components remain stable while allowing task-specific learning in complementary directions. We evaluate REFUSALGUARD across multiple model families, including LLaMA, Gemma, and Qwen, on adversarial safety benchmarks such as AdvBench, DirectHarm4, and JailbreakBench, as well as downstream utility tasks. Our approach achieves attack success rates comparable to base safety-aligned models while maintaining competitive task performance, significantly outperforming baselines.

大模型安全微调拒绝行为几何保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。