arXiv:2505.16104cs.CLcs.CV2025-05ACL被引 3

模型剪枝后安全性能下降?用分层重对齐方法轻量修复。

Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models

  • 按注意力头重要性筛选关键部分,逐层恢复安全相关神经元。
  • 在多个模型和剪枝策略下,安全表现显著提升。
  • 适合关注剪枝后模型安全性的研究者与工程师。

随着大视觉语言模型(LVLMs)规模持续增大,为在资源受限环境部署而采用的网络剪枝技术受到广泛关注。然而我们发现,剪枝常导致安全性能下降。为此,提出一种轻量级新方法——分层安全重对齐(HSR)。HSR首先量化每个注意力头对安全的贡献,识别最关键头,再选择性恢复这些头中对安全起决定作用的神经元,实现从注意力头到神经元层级的安全重构。我们在多种模型和剪枝策略下验证了该方法,均取得显著的安全性能提升。据我们所知,这是首个专注于剪枝后恢复LVLM安全性的研究。

原文摘要 · Abstract (English)

With the increasing size of Large Vision-Language Models (LVLMs), network pruning techniques aimed at compressing models for deployment in resource-constrained environments have garnered significant attention. However, we observe that pruning often leads to a degradation in safety performance. To address this issue, we present a novel and lightweight approach, termed Hierarchical Safety Realignment (HSR). HSR operates by first quantifying the contribution of each attention head to safety, identifying the most critical ones, and then selectively restoring neurons directly within these attention heads that play a pivotal role in maintaining safety. This process hierarchically realigns the safety of pruned LVLMs, progressing from the attention head level to the neuron level. We validate HSR across various models and pruning strategies, consistently achieving notable improvements in safety performance. To our knowledge, this is the first work explicitly focused on restoring safety in LVLMs post-pruning.

模型剪枝安全对齐视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。