对比四种模型去安全化方法,发现数学推理最敏感且效果差异大。
Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
- 用定向正交化技术移除模型拒绝行为,评估四款工具在16个模型上的表现。
- 单次操作方法保留能力更好,数学推理性能变化达-18.81个百分点。
- 适合做安全测试或认知建模的研究者参考工具选型,避免误伤核心能力。
大型语言模型的安全对齐机制通过习得的拒绝行为阻止有害请求,但这也阻碍了认知建模、对抗测试和安全分析等合法研究应用。尽管去安全化技术可通过方向性正交化实现对拒绝表征的精准移除,但现有方法的效果仍缺乏系统评估。本研究在16个指令微调模型(7B-14B参数)上评估了四种工具(Heretic、DECCP、ErisForge、FailSpy),报告了所有模型的兼容性及部分模型的定量指标。单次处理方法在基准子集上展现出更强的能力保留能力(三个模型GSM8K平均变化:ErisForge -0.28 pp;DECCP -0.13 pp);而贝叶斯优化方法导致分布偏移显著(KL散度:0.043–1.646),且能力影响具有模型依赖性。主要发现表明,数学推理能力对去安全化干预最为敏感,其GSM8K得分变化范围为+1.51至-18.81个百分点(相对变化-26.5%),取决于工具选择与模型架构。
原文摘要 · Abstract (English)
Safety alignment mechanisms in large language models prevent responses to harmful queries through learned refusal behavior, yet these same mechanisms impede legitimate research applications including cognitive modeling, adversarial testing, and security analysis. While abliteration techniques enable surgical removal of refusal representations through directional orthogonalization, the relative effectiveness of available implementations remains uncharacterized. This study evaluates four abliteration tools (Heretic, DECCP, ErisForge, FailSpy) across sixteen instruction-tuned models (7B-14B parameters), reporting tool compatibility on all 16 models and quantitative metrics on subsets dictated by tool support. Single-pass methods demonstrated superior capability preservation on the benchmarked subset (avg GSM8K change across three models: ErisForge -0.28 pp; DECCP -0.13 pp), while Bayesian-optimized abliteration produced variable distribution shift (KL divergence: 0.043-1.646) with model-dependent capability impact. These findings provide researchers with evidence-based selection criteria for abliteration tool deployment across diverse model architectures. The principal finding indicates that mathematical reasoning capabilities exhibit the highest sensitivity to abliteration interventions, with GSM8K change ranging from +1.51 pp to -18.81 pp (-26.5% relative) depending on tool selection and model architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。