arXiv:2512.24263cs.AI2025-12

让大模型生成更安全,特别抑制罕见但危险的错误输出

Constrained Language Model Policy Optimization via Risk-aware Stepwise Alignment

  • 用分步对齐法,逐词优化策略以感知风险
  • 显著降低低概率高危害的危险回复,安全性能更强
  • 适合需要高可靠性的实际应用,如医疗或金融对话

在微调预训练语言模型以实现期望行为时,控制风险对保障安全性和可信度至关重要。现有安全对齐方法(如 Safe RLHF 和 SACPO)通常采用风险中性范式,无法有效应对偏离参考策略带来的风险,且对罕见但可能造成灾难性后果的有害行为缺乏鲁棒性。为此,我们提出风险感知分步对齐(RSA),通过引入嵌套风险度量,将安全对齐建模为逐令牌的风险感知约束策略优化问题,并通过分步对齐过程获得基于风险度量的逐令牌策略更新。该设计具有两大优势:(1) 减轻模型过度偏离参考策略引发的风险;(2) 显式抑制低概率但高影响的有害行为。我们在温和假设下提供了策略最优性的理论分析。实验表明,该方法在保持高帮助性的同时,显著增强了安全性,有效抑制了尾部风险,即低概率但高影响的不安全响应。

原文摘要 · Abstract (English)

When fine-tuning pre-trained Language Models (LMs) to exhibit desired behaviors, maintaining control over risk is critical for ensuring both safety and trustworthiness. Most existing safety alignment methods, such as Safe RLHF and SACPO, typically operate under a risk-neutral paradigm that is insufficient to address the risks arising from deviations from the reference policy and offers limited robustness against rare but potentially catastrophic harmful behaviors. To address this limitation, we propose Risk-aware Stepwise Alignment (RSA), a novel alignment method that explicitly incorporates risk awareness into the policy optimization process by leveraging a class of nested risk measures. Specifically, RSA formulates safety alignment as a token-level risk-aware constrained policy optimization problem and solves it through a stepwise alignment procedure that yields token-level policy updates derived from the nested risk measures. This design offers two key benefits: (1) it mitigates risks induced by excessive model shift away from a reference policy, and (2) it explicitly suppresses low-probability yet high-impact harmful behaviors. Moreover, we provide theoretical analysis on policy optimality under mild assumptions. Experimental results demonstrate that our method achieves high levels of helpfulness while ensuring strong safety and significantly suppresses tail risks, namely low-probability yet high-impact unsafe responses.

大模型安全风险控制策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。