让大模型生成更安全、有同理心的回复,无需重训练
ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
- 测试时通过参数高效调节,动态引导生成方向
- 在5个基准上降低有害内容泄露,提升对齐人类价值观
- 适合需要情感敏感响应的客服、医疗等高风险场景
当前语言模型的安全范式在情绪化或高风险场景中表现不足,仅拒绝用户请求可能疏远用户,盲目顺从则加剧风险。我们提出ProSocialAlign,一种测试时、参数高效的框架,无需重训练即可引导生成安全、有同理心且符合价值的回应。我们形式化了五个以人为本的目标,并将安全建模为字典序约束生成:首先应用硬约束排除有害延续;然后在安全集合内优化亲社会质量。该方法结合(i)方向性调控——在参数空间减去学习到的“有害向量”以缓解危害;(ii)偏好感知的自回归奖励建模,联合优化多属性并解决梯度冲突,实现细粒度、用户可控的解码。在五个安全基准上的实证评估显示,该方法达到领先性能,显著减少有害内容泄露,提升对齐人类价值观,多项指标均有明显提升。ProSocialAlign为推理时生成情境敏感、安全且符合人类价值观的回复提供了稳健、模块化的基础。
原文摘要 · Abstract (English)
Current language model safety paradigms often fall short in emotionally charged or high-stakes settings, where refusal-only approaches may alienate users and naive compliance can amplify risk. We propose ProSocialAlign, a test-time, parameter-efficient framework that steers generation toward safe, empathetic, and value-aligned responses without retraining the base model. We formalize five human-centered objectives and cast safety as lexicographic constrained generation: first, applying hard constraints to eliminate harmful continuations; then optimizing for prosocial quality within the safe set. Our method combines (i) directional regulation, a harm-mitigation mechanism that subtracts a learned "harm vector" in parameter space, and (ii) preference-aware autoregressive reward modeling trained jointly across attributes with gradient conflict resolution, enabling fine-grained, user-controllable decoding. Empirical evaluations across five safety benchmarks demonstrate state-of-the-art performance, reducing unsafe leakage and boosting alignment to human values, with strong gains across multiple evaluation metrics. ProSocialAlign offers a robust and modular foundation for generating context-sensitive, safe, and human-aligned responses at inference time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。