从众包偏好中挖掘共性安全准则,让智能体自动遵守。
Implicit Safety Alignment from Crowd Preferences

- 构建分层框架,从众包偏好中提取安全技能
- 无需显式安全奖励,显著降低安全代价
- 适合需要自动安全约束的强化学习场景
基于人类反馈的强化学习(RLHF)可揭示超越任务完成的安全隐含目标。本文聚焦于众包偏好数据集中普遍存在的安全准则:不同用户虽有各异目标,却遵循相似安全原则。我们旨在从众包偏好中发现共享安全准则,并将其迁移至下游强化学习任务,以规范智能体行为、保障安全。实验表明,直接组合奖励的方法存在固有局限;为此,我们提出基于安全众包偏好的强化学习框架,通过高层策略组合从偏好中提取的安全技能,实现安全解题。在多个安全强化学习环境及一个具有多样化用户目标和共同安全约束的类大模型任务上,该方法在无显式安全奖励条件下显著降低安全成本,同时任务表现接近使用真实安全信号训练的最优方法。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) can reveal implicit objectives such as safety considerations that go beyond task completion. In this work, we focus on the common safety criteria embedded in crowd preference datasets, where different users may express distinct preferences or objectives, yet follow similar safety principles. Our aim is to discover shared safety criteria from crowd preferences and then transfer them to downstream RL tasks to regularize agent behavior and enforce safety. We first show that direct reward combination-optimizing a preference-learned reward model together with downstream task rewards-has inherent limitations. Motivated by this, we propose Safe Crowd Preference-based RL, a hierarchical framework that extracts safety-aligned skills from crowd preferences and composes them via a high-level policy to safely solve downstream tasks. Experiments across safe RL environments and a preliminary LLM-style task with diverse user goals and shared safety constraints demonstrate that our approach substantially lowers safety costs without access to explicit safety rewards, while achieving task performance comparable to oracle methods trained with ground-truth safety signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。