通过自适应多方向权重调整,精准修改大模型行为而不损整体性能。
Gabliteration: Adaptive Multi-Directional Neural Weight Modification for Selective Behavioral Alteration in Large Language Models
- 采用自适应多方向投影与正则化层选择,动态优化权重修改路径。
- 在0.6B至4B参数模型上验证,实现行为修改且无关领域质量下降小。
- 适合需要定制化模型行为但又不想牺牲通用能力的研究者使用。
我们提出Gabliteration,一种新型神经权重修改技术,突破传统消融方法的局限,通过自适应多方向投影与正则化层选择,解决现有方法在修改特定行为时损害模型整体质量的问题。借助动态层优化、正则化投影矩阵和自适应缩放机制,实现理论上更优的权重修改,同时最小化对无关领域的性能影响。我们在Hugging Face上发布了gabliterated-v1系列模型(0.6B至4B参数),验证了该方法在多种模型规模下的实用性。
原文摘要 · Abstract (English)
We present Gabliteration, a novel neural weight modification technique that advances beyond traditional abliteration methods by implementing adaptive multi-directional projections with regularized layer selection. Our approach addresses the fundamental limitation of existing methods that compromise model quality while attempting to modify specific behavioral patterns. Through dynamic layer optimization, regularized projection matrices, and adaptive scaling mechanisms, we achieve theoretically superior weight modification while minimizing quality degradation in unrelated domains. We validate our method through the gabliterated-v1 model series (0.6B to 4B parameters) available on Hugging Face, demonstrating practical applicability across multiple model scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。