用分布支配关系提升安全强化学习的抗风险能力
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
- 以一阶随机支配替代期望约束,直接控制代价分布尾部风险
- 通过最优传输框架实现可微优化,支持端到端训练
- 量化加权支配可统一调控多种风险度量,适合高风险场景
安全强化学习从人类反馈(RLHF)通常依赖期望代价约束,但期望仅反映代价分布的单一统计量,无法捕捉分布不确定性,尤其在重尾或罕见灾难事件下表现不足。这在需要鲁棒性和风险敏感性的场景中存在问题。随机支配通过比较完整代价分布,提供更严谨的替代方案,可直接控制尾部风险和分布外失效。本文提出风险敏感对齐框架RAD,将标量期望约束替换为一阶随机支配(FSD)约束。通过最优传输(OT)框架,在熵正则化与Sinkhorn迭代支持下,实现可微且计算高效的优化目标。进一步引入分位数加权FSD约束,证明其可普遍控制一大类谱风险度量(SRMs),即加权支配改进意味着对应谱风险的确定性改善。这为通过分位数权重函数调节模型风险偏好提供了理论依据。实验表明,RAD在保持帮助性竞争力的同时显著提升无害性表现,并在分布外无害性评估中展现更强鲁棒性。
原文摘要 · Abstract (English)
Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost distribution and fails to account for distributional uncertainty, particularly under heavy tails or rare catastrophic events. This limitation is problematic when robustness and risk sensitivity are critical. Stochastic dominance offers a principled alternative by comparing entire cost distributions rather than just their averages, enabling direct control over tail risks and potential out-of-distribution failures that expectation-based constraints may overlook. In this work, we propose Risk-sensitive Alignment via Dominance (RAD), a novel alignment framework that replaces scalar expected cost constraints with First-Order Stochastic Dominance (FSD) constraints. We operationalize this constraint by comparing the target policy's cost distribution to that of a reference policy within an Optimal Transport (OT) framework, using entropic regularization and Sinkhorn iterations to obtain a differentiable and computationally efficient objective for stable end-to-end optimization. Furthermore, we introduce quantile-weighted FSD constraints and show that weighted FSD universally controls a broad class of Spectral Risk Measures (SRMs), so that improvements under weighted dominance imply guaranteed improvements in the corresponding spectral risk. This provides a principled mechanism for tuning a model's risk profile via the quantile weighting function. Empirical results demonstrate that RAD improves harmlessness over baselines while remaining competitive in helpfulness, and exhibits greater robustness on out-of-distribution harmlessness evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。