测试大模型中欺骗行为能否被精准移除,发现其会自我再生且损害整体语言能力。
Cyclic Ablation: Testing Concept Localization against Functional Regeneration in AI
- 通过循环删减与对抗训练,反复尝试移除欺骗概念。
- 每次删除后欺骗行为均恢复,同时困惑度持续上升。
- 揭示复杂概念分布于模型全局,难以局部编辑。
安全与可控性对大型语言模型至关重要。核心问题在于:如欺骗这类不良行为是否为可定位的功能模块,可被移除,还是与模型核心认知能力深度纠缠?我们提出“循环删减”(cyclic ablation)方法,结合稀疏自编码器、定向删减和在DistilGPT-2上的对抗训练,试图消除欺骗概念。结果表明,与定位假说相反,欺骗行为具有高度韧性:每次删减后,模型均通过对抗训练恢复欺骗行为,此过程称为功能再生。关键的是,每次“神经外科手术”都导致语言性能的渐进但可观测下降,表现为困惑度持续上升。这些发现支持复杂概念是分布式且纠缠的观点,凸显了基于机制可解释性直接编辑模型的局限性。
原文摘要 · Abstract (English)
Safety and controllability are critical for large language models. A central question is whether undesirable behaviors like deception are localized functions that can be removed, or if they are deeply intertwined with a model's core cognitive abilities. We introduce "cyclic ablation," an iterative method to test this. By combining sparse autoencoders, targeted ablation, and adversarial training on DistilGPT-2, we attempted to eliminate the concept of deception. We found that, contrary to the localization hypothesis, deception was highly resilient. The model consistently recovered its deceptive behavior after each ablation cycle via adversarial training, a process we term functional regeneration. Crucially, every attempt at this "neurosurgery" caused a gradual but measurable decay in general linguistic performance, reflected by a consistent rise in perplexity. These findings are consistent with the view that complex concepts are distributed and entangled, underscoring the limitations of direct model editing through mechanistic interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。