arXiv:2511.00029cs.LGcs.AI2025-11被引 1

用稀疏自编码器精准控制大模型拒绝率,兼顾安全与可用性。

Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts

  • 通过对比提示词筛选最优特征,实现对大模型的定向引导。
  • 在Llama-3 8B上提升18.9%安全性,同时提高11.1%实用性。
  • 适合关注大模型安全可控部署的研究者与工程师。

大语言模型部署需使其识别并拒绝不安全提示,同时响应安全提示。现有方法需调整模型权重等高成本操作。尽管稀疏自编码器(SAEs)已实现对大模型特征的可解释提取,但缺乏系统化的特征选择方法和对安全-效用权衡的严谨评估。本文探索使用不同引导特征与强度的SAE引导策略,结合来自teknium/OpenHermes-2p5-Mistral-7B和Air Bench eu-dataset的AI生成提示数据集,提出一种高效选取模型最佳引导特征的对比提示方法。在Llama-3 8B上测试表明,该方法在实现18.9%安全性能提升的同时,还使模型效用增加11.1%,证明通过有原则的特征选择,目标导向的SAE引导可突破传统安全-效用权衡瓶颈。

原文摘要 · Abstract (English)

Large Language Model (LLM) deployment requires guiding the LLM to recognize and not answer unsafe prompts while complying with safe prompts. Previous methods for achieving this require adjusting model weights along with other expensive procedures. While recent advances in Sparse Autoencoders (SAEs) have enabled interpretable feature extraction from LLMs, existing approaches lack systematic feature selection methods and principled evaluation of safety-utility tradeoffs. We explored using different steering features and steering strengths using Sparse Auto Encoders (SAEs) to provide a solution. Using an accurate and innovative contrasting prompt method with the AI-Generated Prompts Dataset from teknium/OpenHermes-2p5-Mistral-7B and Air Bench eu-dataset to efficiently choose the best features in the model to steer, we tested this method on Llama-3 8B. We conclude that using this method, our approach achieves an 18.9% improvement in safety performance while simultaneously increasing utility by 11.1%, demonstrating that targeted SAE steering can overcome traditional safety-utility tradeoffs when optimal features are identified through principled selection methods.

大模型安全稀疏编码器提示工程效用优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。