提出新方法让专家模型更精准地绕过拒绝指令
Expert-Aware Refusal Steering

- 利用特定专家路由模式与方向,精准控制拒绝行为
- 仅凭一个专家输出即可有效诱导模型回复有害请求
- 揭示注意力机制在专家模型拒绝行为中的关键作用
指令微调的大语言模型(LLMs)安全对齐依赖于其可靠拒绝有害或不允许请求的能力。近期研究发现,在推理阶段施加引导向量可有效抑制拒绝行为,使模型响应有害请求。本文将该拒绝引导方法扩展至三个开源混合专家(MoE)LLM,并发现其性能不受MoE架构复杂路由模式影响。进一步提出两种基于专家的拒绝引导方法,利用拒绝相关的专家路由模式与专家特异性引导方向,抑制正常拒绝行为。结果表明,仅依据单个专家输出即可有效引导拒绝行为。研究发现,引导方法捕捉到的拒绝信号与专家路由行为存在差异,表明注意力机制在MoE模型拒绝行为中起着重要作用。
原文摘要 · Abstract (English)
Safety alignment in instruction-tuned large language models (LLMs) depends on a model's ability to reliably refuse to respond to harmful or disallowed requests. Recent work has shown that a steering vector can be applied to a dense LLM during inference to effectively suppress refusal behavior, inducing response to harmful requests. We extend this refusal steering method to three open-source Mixture-of-Experts (MoE) LLMs and find that steering performance is uninhibited by the complex routing patterns inherent to the MoE architecture. We then propose two expert-aware refusal steering methods that leverage refusal-specific expert routing patterns and expert-specific steering directions to suppress normal refusal behavior. We find that refusal behavior can be effectively steered based on the output of a single expert. Our results show that refusal signals captured by steering methods differ from expert routing behavior, suggesting a substantial role for attention in MoE refusal behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。