无需输出提示,自动识别大模型拒绝行为的隐藏方向。
COSMIC: Generalized Refusal Direction Identification in LLM Activations
- 用余弦相似度自动找可操控的神经激活方向。
- 在对抗和弱对齐模型中仍能准确识别拒绝方向。
- 适合研究模型安全与可控性的研究人员使用。
大型语言模型(LLMs)在其激活空间中编码了诸如拒绝等行为,但识别这些行为仍具挑战性。现有方法通常依赖预定义的拒绝模板或需人工分析。我们提出COSMIC(Cosine Similarity Metrics for Inversion of Concepts),一种完全独立于模型输出的自动化方向选择框架,通过余弦相似度识别可行的控制方向及目标层。COSMIC在无需假设模型存在特定拒绝标记的情况下,达到与先前方法相当的控制性能。它在对抗性设置和弱对齐模型中均能可靠识别拒绝方向,并以极小的误拒增加将模型引导至更安全行为,展现出在多种对齐条件下的鲁棒性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) encode behaviors such as refusal within their activation space, yet identifying these behaviors remains a significant challenge. Existing methods often rely on predefined refusal templates detectable in output tokens or require manual analysis. We introduce \textbf{COSMIC} (Cosine Similarity Metrics for Inversion of Concepts), an automated framework for direction selection that identifies viable steering directions and target layers using cosine similarity - entirely independent of model outputs. COSMIC achieves steering performance comparable to prior methods without requiring assumptions about a model's refusal behavior, such as the presence of specific refusal tokens. It reliably identifies refusal directions in adversarial settings and weakly aligned models, and is capable of steering such models toward safer behavior with minimal increase in false refusals, demonstrating robustness across a wide range of alignment conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。