通过精细调控神经元实现大模型安全与性能的平衡
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
- 按攻击感知定位关键神经元,动态调节其激活
- 在保持文本质量的同时显著提升抗攻击能力
- 适合需要高安全与强可用性的大模型部署场景
确保大语言模型在可靠部署中兼具鲁棒安全性与实用性能至关重要。现有方法普遍存在双重缺陷:对恶意攻击防御不足、对正常请求频繁拒绝,同时导致生成文本质量下降及通用任务性能退化,前者体现为安全与鲁棒性的矛盾,后者反映实用性受损。我们归因于现有方法采用粗粒度的层级干预。为此,提出NeuronTune——一种细粒度神经元调制框架,通过攻击感知归因精准定位各层中的安全关键与实用保留神经元,并利用元学习自适应调节其激活值。关键的是,该方法通过神经元包含阈值动态控制干预范围,灵活支持以安全性优先或实用性优先为目标。大量实验表明,该方法优于当前最先进方案,在保障卓越安全性的同时维持优秀实用性。
原文摘要 · Abstract (English)
Ensuring robust safety alignment while preserving utility is critical for the reliable deployment of Large Language Models (LLMs). However, current techniques fundamentally suffer from intertwined deficiencies: insufficient robustness against malicious attacks, frequent refusal of benign queries, degradation in generated text quality and general task performance, the former two reflecting tensions in robust safety and the latter constituting utility impairment. We attribute these limitations to the coarse-grained layer-wise interventions in existing methods. To resolve this, we propose NeuronTune, a fine-grained framework that pinpoints and modulates sparse neurons to achieve simultaneous safety-utility optimization. Our approach first pinpoints safety-critical and utility-preserving neurons across all layers via attack-aware attribution, then adapts meta-learning to adaptively modulate their activations. Crucially, the intervention scope of NeuronTune is dynamically controlled via neuron inclusion thresholds, providing a flexible mechanism to prioritize either security-critical or utility-priority requirements. Extensive experimental results demonstrate that our method outperforms existing state-of-the-art technologies, achieving superior model safety while maintaining excellent utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。