arXiv:2602.17696cs.LGcs.AI2026-02ACL

现有方法难以稳定识别通用安全参数区域。

Can LLM Safety Be Ensured by Constraining Parameter Regions?

  • 测试四种不同粒度的参数区域定位方法
  • 安全区域重叠度低,且随效用数据细化进一步下降
  • 适用于研究安全机制的可迁移性与局限性

大型语言模型(LLMs)常被认为存在“安全区域”——即修改某些参数可直接影响其安全行为。本文系统评估了四种跨不同参数粒度(从单个权重到整个Transformer层)的安全区域识别方法,在四个不同规模的骨干模型家族中进行测试。基于十个安全识别数据集,发现所识别的安全区域交集率(IoU)仅为低至中等水平。当使用非有害查询的效用数据进一步精炼这些区域时,重叠度显著下降。结果表明,当前技术无法可靠识别出稳定、数据无关的安全参数区域。

原文摘要 · Abstract (English)

Large language models (LLMs) are often assumed to contain ``safety regions'' -- parameter subsets whose modification directly influences safety behaviors. We conduct a systematic evaluation of four safety region identification methods spanning different parameter granularities, from individual weights to entire Transformer layers, across four families of backbone LLMs with varying sizes. Using ten safety identification datasets, we find that the identified safety regions exhibit only low to moderate overlap, as measured by IoU. The overlap drops significantly when the safety regions are further refined using utility datasets (\ie non-harmful queries). These results suggest that current techniques fail to reliably identify a stable, dataset-agnostic safety region.

安全区域参数约束大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。