arXiv:2607.02714cs.CRcs.AI2026-07

研究发现大模型安全机制会阻碍合法网络安全操作,提出针对性关闭策略。

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale

论文配图:Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale
图 1 · 摘自论文原文
  • 通过消融实验定位网络安全相关拒绝响应的分布特征
  • 在24个开源大模型上验证了1万亿参数模型可实现领域特异性关闭
  • 按关闭难易度将模型分为3类,揭示训练方式与架构的关键影响

安全对齐是大语言模型训练中的必要步骤,但其概念未区分不同领域的潜在危害程度,这在网络安全领域引发严重问题——模型不应因安全机制限制而无法执行合法授权操作。本文基于对24个开源大模型的大规模消融实验,以1万亿参数的Kimi K2为例,证明了领域特异性消融的可行性。结合近期发现的拒绝行为在模型层间呈多维子空间分布的研究,我们进一步发现该现象在万亿参数MoE架构中广泛存在。本研究聚焦于仅包含网络安全有害概念的拒绝部分,分析模型特征与消融效果的相关性,识别出安全训练类型和模型架构是最可靠的预测因子。最终,我们将模型划分为3个消融敏感度等级,并提出若干假设,解释特定干预在不同模型中产生差异的原因。

原文摘要 · Abstract (English)

There is no doubt that safety alignment is an essential step in LLM training. However, conceptually it does not distinguish between various domains and the level of potential harm of a query, which creates significant complications in the fields like cyber security, where a model should not be constrained by its safety circuits to accomplish the goals of legitimate, authorized operations. In this work, we share our findings from a large scale abliteration experiment on 24 open-source LLMs and show that domain-specific abliteration is achievable with standard methodology on the example of a 1T-parameter Kimi K2. Building on recent work showing that refusal in LLMs occupies a multi-dimensional subspace within layers, we find that it is also distributed widely across layers, especially in trillion-parameter MoE architectures, and so we aim to capture the part of it that represents harmful concepts in the cybersecurity domain exclusively. We also investigate the correlation between models' features and the effect of domain-specific abliteration, identifying that the type of safety training and architecture are the most reliable predictors. Finally, we classify the models into 3 abliteration susceptibility tiers and put forward a set of conjectures as to why a particular effect from this intervention might be observed in a given model.

大模型安全消融实验网络安全模型可控性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。