让AI模型能灵活适应不同安全规则,自动调整响应策略。
Configurable Reward Model for Balanced Safety Alignment

- 通过可配置奖励模型实现安全规则动态调整
- 在新安全配置上达到94.6%和75.8%的准确率
- 无需人工标注,适合需要多场景安全适配的系统
将大语言模型(LLMs)对异构且快速变化的安全需求进行对齐仍是重大挑战。现有指令微调模型和独立安全分类器难以泛化到新的安全配置,因此需要能够显式配置以适应变化规范的奖励模型(RMs)。本文提出可配置安全奖励模型(CSRM),在校准安全性合规性和奖励建模上联合优化。方法采用针对配置的数据增强,确保指令遵循同时保持相对严重性结构。所得到的模型对细粒度安全配置和对话细微差别敏感,显著提升对未见过的安全配置的泛化能力。CSRM在近期可配置安全基准测试中表现领先,包括CoSApien(94.6% F1)和DynaBench(75.8% F1),且无需额外人工标注。用于下游安全对齐时,相比现有基线,使模型在帮助性与安全性之间取得更优平衡。
原文摘要 · Abstract (English)
Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned LLMs and standalone safety classifiers often fail to generalize to new safety configurations, motivating the need for Reward Models (RMs) that are explicitly configurable to changing specifications. We introduce the Configurable Safety Reward Model (CSRM), which is jointly optimized for calibrated safety compliance and reward modeling. Our approach is supported by configuration-targeted data augmentation that enforces instruction adherence while preserving relative severity structure. The resulting RM is sensitive to fine-grained safety configurations and conversational nuances, substantially improving generalization to previously unseen safety configurations. CSRM achieves state-of-the-art performance on recent configurable safety benchmarks, including CoSApien (94.6% F1) and DynaBench (75.8% F1), without requiring additional human annotation. When used for downstream safety alignment, CSRM yields LLMs with a significantly improved helpfulness-safety tradeoff compared to existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。