通过分离知识原子实现大模型行为精准控制,提升安全性和鲁棒性。
Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
- 提出STA方法,用稀疏自编码器定位可独立操控的知识原子。
- 在对抗场景下验证,控制精度高且无意外副作用。
- 适用于大模型推理控制,适合安全敏感应用研究者。
精确控制语言模型生成对保障其安全与可靠性至关重要。尽管提示工程和引导调控常用于干预模型行为,但模型参数量庞大导致内部表征高度耦合,限制了控制精度并可能引发意外后果。近期研究尝试使用稀疏自编码器(SAE)在高维空间中解耦知识以实现引导,但因难以定位原子级知识组件,应用仅限于简单任务。本文提出新的Steering Target Atoms(STA)方法,通过隔离并操控解耦的知识组件来增强安全性。全面实验表明该方法有效,进一步分析显示其在对抗场景中具有更强鲁棒性与灵活性。我们还将该策略应用于大型推理模型,证实其在精准推理控制中的有效性。
原文摘要 · Abstract (English)
Precise control over language model generation is vital for ensuring both safety and reliability. Although prompt engineering and steering are commonly used to intervene in model behaviors, the vast number of parameters in models often results in highly intertwined internal representations. This interdependency can limit control precision and sometimes lead to unintended side effects. Recent research has explored the use of sparse autoencoders (SAE) to disentangle knowledge in high-dimensional spaces for steering. However, these applications have been limited to toy tasks owing to the nontrivial issue of locating atomic knowledge components. In this paper, we propose Steering Target Atoms (STA), a novel method that isolates and manipulates disentangled knowledge components to enhance safety. Comprehensive experiments demonstrate the effectiveness of our approach. Further analysis reveals that steering exhibits superior robustness and flexibility, particularly in adversarial scenarios. We also apply the steering strategy to the large reasoning model, confirming its effectiveness in precise reasoning control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。