arXiv:2410.01174cs.CLcs.AI2024-10被引 21

通过类别特异性向量实现大模型推理时的安全控制

Towards Inference-time Category-wise Safety Steering for Large Language Models

  • 用特定类别的向量在推理阶段精准调节模型输出
  • 提取有效向量提升安全控制效果,同时保持生成质量
  • 无需训练,适合对安全性要求高的实际应用

尽管大语言模型在各类应用场景中取得了前所未有的进展,其安全对齐仍是活跃研究领域。即使经过大量对齐与安全训练,大模型仍存在脆弱性,需通过无需训练的推理阶段方法进行额外安全引导。已有工作探索了潜在表示空间中的激活如何编码概念,并通过表征工程诱导模型输出特定概念,但将其应用于安全领域的研究相对不足。本文提出一种新的推理阶段安全引导方法:(i) 使用类别特异性引导向量,实现更精细的控制;(ii) 采用先进的向量提取技术,提升引导效果并保持生成文本质量。我们在多个大模型和数据集上验证该方法的有效性,并讨论其影响与最佳实践。

原文摘要 · Abstract (English)

While large language models (LLMs) have seen unprecedented advancements in capabilities and applications across a variety of use-cases, safety alignment of these models is still an area of active research. The fragile nature of LLMs, even models that have undergone extensive alignment and safety training regimes, warrants additional safety steering steps via training-free, inference-time methods. While recent work in the area of mechanistic interpretability has investigated how activations in latent representation spaces may encode concepts, and thereafter performed representation engineering to induce such concepts in LLM outputs, the applicability of such for safety is relatively under-explored. Unlike recent inference-time safety steering works, in this paper we explore safety steering of LLM outputs using: (i) category-specific steering vectors, thereby enabling fine-grained control over the steering, and (ii) sophisticated methods for extracting informative steering vectors for more effective safety steering while retaining quality of the generated text. We demonstrate our exploration on multiple LLMs and datasets, and showcase the effectiveness of the proposed steering method, along with a discussion on the implications and best practices.

大模型安全推理阶段向量引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。