用可解释概念识别模型漏洞,构建无需微调的智能安全防护
ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
- 通过稀疏自编码器挖掘模型内部与越狱攻击相关的可解释概念
- 发现越狱攻击在表征空间中存在共享激活几何结构,提升防御通用性
- 无需微调即可实现可解释、强鲁棒的安全防护,适合高安全性场景
大语言模型在各类应用中取得成功,但其安全性仍受越狱方法威胁。尽管对齐与安全微调提供一定防护,仍难以抵御隐蔽诱导生成有害内容的攻击,易引发针对性滥用和用户画像泄露。本文提出ConceptGuard框架,利用稀疏自编码器(SAEs)识别大模型内部与不同越狱主题相关联的可解释概念。通过提取语义明确的内部表征,ConceptGuard构建出无需微调、完全可解释且具备泛化能力的安全防护机制。基于大模型机械可解释性的进展,该方法揭示越狱攻击在表征空间中存在共享激活几何结构,为设计更可解释、更通用的防御策略提供了可能基础。
原文摘要 · Abstract (English)
Large Language Models have found success in a variety of applications. However, their safety remains a concern due to the existence of various jailbreaking methods. Despite significant efforts, alignment and safety fine-tuning only provide a certain degree of robustness against jailbreak attacks that covertly mislead LLMs towards the generation of harmful content. This leaves them prone to a range of vulnerabilities, including targeted misuse and accidental user profiling. This work introduces \textbf{ConceptGuard}, a novel framework that leverages Sparse Autoencoders (SAEs) to identify interpretable concepts within LLM internals associated with different jailbreak themes. By extracting semantically meaningful internal representations, ConceptGuard enables building robust safety guardrails -- offering fully explainable and generalizable defenses without sacrificing model capabilities or requiring further fine-tuning. Leveraging advances in the mechanistic interpretability of LLMs, our approach provides evidence for a shared activation geometry for jailbreak attacks in the representation space, a potential foundation for designing more interpretable and generalizable safeguards against attackers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。