arXiv:2607.05355cs.CLcs.ET2026-07被引 1

用直接因果测试发现:越能稳定选中神经元的模型,反而越不靠谱。

Faithfulness to Refusal: A Causal Audit of Neuron Selectors

论文配图:Faithfulness to Refusal: A Causal Audit of Neuron Selectors
图 1 · 摘自论文原文
  • 通过一次性置零神经元行,直接检验不同方法选出的关键神经元是否真有因果作用。
  • 基于归因的神经元可有效触发拒绝有害内容,且保持语言流畅性,随机对照组则失败。
  • 不同方法选出的神经元集几乎不重叠,说明拒绝机制具有冗余性,非唯一路径。

归因分数被广泛用于识别语言模型中对剪枝、可解释性和安全编辑重要的神经元行,但其是否真正反映因果重要性尚未直接验证。本文构建了两个基于单次神经元行置零的配对审计:首先在语言建模层面评估,归因方法显著优于激活值和幅度基线,在五种大型语言模型中准确识别出可移除的神经元行;随后将同一干预转化为行为测试,利用对比有害与无害信号,发现归因选出的神经元足以实现对仇恨和犯罪内容的拒绝,同时避免过度拒绝良性内容,并保持语言流畅性;而同层深度的随机对照组则无法实现。高排名稳定性选择器反而可能因果有效性最低。拒绝行为存在于冗余子空间,不同归因方法通过几乎不重叠的神经元集实现拒绝,表明恢复的编辑只是充分条件集合的一种实现,而非唯一机制。这些发现表明,排名稳定性无法捕捉直接因果审计所揭示的选择器失效问题。

原文摘要 · Abstract (English)

Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing for safety, yet whether they identify causally important rows is rarely tested directly. We address this with two paired audits built on one-shot neuron-row zeroing. We first audit selectors at the language-modeling level: attribution methods substantially outperform activation and magnitude-based baselines at identifying dispensable rows across five LLMs. We then adapt the same intervention into a behavior test by driving it with a contrastive harmful-versus-benign signal; the attributed rows are sufficient to install refusal on hate and crime while keeping benign over-refusal low and preserving language model fluency, and specific in that layer-matched random controls at the same depths fail. Highly rank-stable selectors can be among the least causally valid. Refusal moreover lives in a redundant subspace, where different attribution methods install it through largely disjoint row sets, so the recovered edit is one realization of a sufficient set rather than a unique mechanism. Together, these findings show that rank-stability proxies miss the kinds of selector failures a direct causal audit can surface.%

可解释性因果审计神经元选择语言模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。