arXiv:2410.08968cs.CLcs.AI2024-10ICLR被引 42

让大模型按需调整安全标准,无需重训练即可适应不同文化与用户需求。

Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements

  • 通过自然语言描述安全配置,在推理时动态调整模型行为。
  • 相比基线方法,可控性显著提升,帮助性与安全配置一致性更高。
  • 适合需要灵活适配多地区、多用户安全要求的场景。

当前大模型的安全对齐普遍采用一刀切策略:模型拒绝所有被厂商认定为不安全的内容。这在跨文化背景下缺乏灵活性,且无法满足用户多样化安全需求,静态安全标准既限制实用性又增加重对齐成本。本文提出可控安全对齐(CoSA)框架,无需重训练即可根据用户需求动态调整模型安全行为。通过将安全配置(即自由形式的自然语言描述)作为系统提示的一部分,授权用户可在推理阶段修改配置以调节模型行为。为此,我们提出 CoSAlign 方法,一种数据驱动的对齐机制,使模型能快速响应不同安全配置。我们设计了新的可控性评估协议,综合衡量帮助性与配置安全,得出 CoSA-Score;并构建 CoSApien 基准,包含真实应用场景下的多样安全需求及对应评测提示。实验表明,CoSAlign 在可控性上显著优于强基线,包括上下文对齐方法。该框架推动大模型更好地体现多元人类价值观,提升实际应用价值。

原文摘要 · Abstract (English)

The current paradigm for safety alignment of large language models (LLMs) follows a one-size-fits-all approach: the model refuses to interact with any content deemed unsafe by the model provider. This approach lacks flexibility in the face of varying social norms across cultures and regions. In addition, users may have diverse safety needs, making a model with static safety standards too restrictive to be useful, as well as too costly to be re-aligned. We propose Controllable Safety Alignment (CoSA), a framework designed to adapt models to diverse safety requirements without re-training. Instead of aligning a fixed model, we align models to follow safety configs -- free-form natural language descriptions of the desired safety behaviors -- that are provided as part of the system prompt. To adjust model safety behavior, authorized users only need to modify such safety configs at inference time. To enable that, we propose CoSAlign, a data-centric method for aligning LLMs to easily adapt to diverse safety configs. Furthermore, we devise a novel controllability evaluation protocol that considers both helpfulness and configured safety, summarizing them into CoSA-Score, and construct CoSApien, a human-authored benchmark that consists of real-world LLM use cases with diverse safety requirements and corresponding evaluation prompts. We show that CoSAlign leads to substantial gains of controllability over strong baselines including in-context alignment. Our framework encourages better representation and adaptation to pluralistic human values in LLMs, and thereby increasing their practicality.

安全对齐推理时调整可控性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。