arXiv:2607.17427cs.LGcs.CL2026-07

删除拒绝指令会改变模型决策风格,导致乐观度上升和自信表达变化。

Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal on Decision Disposition Across Model Families

  • 通过删去拒绝方向权重,测试模型在不确定决策中的表现变化。
  • 删减后模型更乐观(Gemma+12.2pp,Qwen+7.4pp),自我解释更长,不确定性词汇减少。
  • 同一操作对不同模型产生相反的自信影响,说明‘去审查’并非简单移除功能。

删除模型拒绝方向权重是当前主流‘去审查’开源模型的标准做法。我们发现这种操作并不干净:在60只华沙证券交易所股票、18周共21,600次涨跌决策任务中,使用冻结推理管道对比基线与删减版本的两个Mixture-of-Experts模型(Gemma-4-26B-A4B-it 和 Qwen3-30B-A3B-Instruct-2507)。任务不触发拒绝行为,因此所有差异均为副作用。结果显示:删减模型系统性更乐观(Gemma+12.2pp,Qwen+7.4pp),自解释更冗长,强制自我批判中使用更少明确不确定性词汇;而自信程度变化反向:Gemma删减后更不自信,Qwen则更自信,且置信区间不重叠。能力相关变量排除了指令遵循退化的可能,且无模型具备经济收益能力。此外,工具链中存在两个独立污染源:量化器错配与过时对话模板,暗示社区修改检查点研究普遍受工具链缺陷影响。部署‘去审查’模型即等同于部署一个决策行为被重构的实体,而非仅去除拒绝功能的原模型。

原文摘要 · Abstract (English)

Abliteration - deleting a model's refusal direction from its weights - is the standard recipe behind popular "uncensored" open-weight models. We show the surgery is not clean. As a disposition probe we use 21,600 decisions under uncertainty - weekly up/down calls on 60 Warsaw Stock Exchange equities over 18 weeks, replayed through a frozen pipeline so the decision-layer model is the only variable. The task elicits no refusals at all, so any between-arm delta is pure side effect. Holding provenance constant (official BF16 checkpoints, a single abliteration author, an identical serving stack, one byte-identical frozen prompt), we compare base and abliterated arms of two Mixture-of-Experts families, Gemma-4-26B-A4B-it and Qwen3-30B-A3B-Instruct-2507. Three effects replicate across both families (weeks-clustered bootstrap CIs excluding zero): abliterated models are systematically more optimistic (+12.2 pp Gemma, +7.4 pp Qwen; the confirmed preregistered endpoint), justify themselves at greater length, and use fewer explicit uncertainty words in forced self-critiques (both exploratory). A fourth effect reverses sign: the same operation makes Gemma-abliterated less confident and Qwen-abliterated more (family CIs non-overlapping) - one weight surgery, opposite shifts in expressed confidence. Capability covariates rule out instruction-following degradation as the driver, and no arm shows economic skill: the apparent edge of abliterated arms is regime beta, not alpha. Our provenance audit also caught two independent contamination channels - a mismatched-quantizer pilot pair and a stale community chat template that silently mangled the rendered prompt - suggesting toolchain artifacts are the rule in studies of community-modified checkpoints. Whoever deploys an "uncensored" model as an agent is deploying a measurably different decision-maker, not the base model minus refusals.

模型安全去审查决策偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。