arXiv:2410.12194cs.CL2024-10被引 3

用负面提示引导模型避开有害输出,提升对齐人类价值观的能力

Negative-Prompt-driven Alignment for Generative Language Model

  • 引入负面提示生成不良响应,与正面样本一起优化模型
  • 在多个数据集上显著降低有害内容生成率,提升价值对齐效果
  • 适合需要避免风险输出的场景,如医疗、法律等安全敏感领域

大语言模型虽具强大能力,但其输出与人类价值观对齐仍面临挑战。现有对齐方法多依赖正向示例,忽视负面响应在引导模型规避不良行为中的作用。例如,广泛使用的对齐数据集缺乏明确违背人类价值观的负面样本,导致训练中难以抑制有害或偏见输出。为此,本文提出NEAT(负提示驱动对齐),通过在优化过程中引入负提示生成不良响应,显式惩罚模型产生有害内容的行为。该双重反馈机制不仅引导模型趋向理想输出,还主动规避偏差与有害响应。基于预训练模型,NEAT采用在线对齐策略,结合包含正负样本的扩展偏好数据集,使用排名损失进行训练。大量实验证明,NEAT能显著提升模型与人类价值观和偏好的对齐程度。

原文摘要 · Abstract (English)

Large language models have achieved remarkable capabilities, but aligning their outputs with human values and preferences remains a significant challenge. Existing alignment methods primarily focus on positive examples while overlooking the importance of negative responses in guiding models away from undesirable behaviors. For instance, the widely-used alignment datasets reveals a scarcity of explicit negative examples that contradict human values, hindering its ability to discourage harmful or biased outputs during training. To address this limitation, we propose NEAT, i.e., NEgative-prompt-driven AlignmenT, to introduce negative prompts to generate undesirable responses alongside positive examples during the optimization process. NEAT explicitly penalizes the model for producing harmful outputs, guiding it not only toward desirable behaviors but also steering it away from generating undesirable, biased responses. This dual feedback mechanism enables better alignment with human preferences, crucial in contexts where avoiding harm is paramount. Starting from a pre-trained language model, NEAT performs online alignment by incorporating a ranking loss derived from an expanded preference dataset containing both positive and negative examples. Extensive experiments validate NEAT's effectiveness in significantly enhancing language models' alignment with human values and preferences.

语言模型价值对齐负提示安全生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。