系统梳理扩散模型对齐与安全的强化学习方法,揭示关键挑战与技术路径。
Alignment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey
- 从人类反馈到奖励建模,多种优化机制实现文本生成对齐。
- 对比不同方法在计算成本、抗攻击性与安全性上的表现差异。
- 适合关注生成模型安全、对齐与可信赖部署的研究者阅读。
扩散模型已成为图像与多模态生成的核心范式,但其部署仍面临对齐性、安全性、偏好满足及滥用鲁棒性等持续问题。本文综述了通过强化学习、奖励建模、偏好优化和安全特化微调实现文本到图像扩散模型对齐的最新进展。从反馈来源、奖励信号形式、优化机制、分布偏移与奖励过优化处理,以及安全是否作为显式约束五个维度组织文献。涵盖人类反馈强化学习、KL正则化策略优化、直接偏好优化、二元效用优化、可微奖励微调、代理奖励学习、区域感知微调及安全导向的DPO变体。为提升可读性,附有扩散采样、奖励建模与偏好优化的教程说明,并简要关联图像扩散对齐与新兴文本及掩码语言扩散模型。比较代表性方法在反馈需求、计算开销、可扩展性、抗奖励劫持能力及安全关键部署适用性方面的差异。最后归纳出五大开放挑战:多目标对齐、高效反馈偏好学习、对抗鲁棒的安全对齐、动态规范下的持续对齐、可解释的奖励建模。旨在为扩散模型对齐领域提供技术地图,并指明可靠部署前必须解决的方法缺口。
原文摘要 · Abstract (English)
Diffusion models have become a central paradigm for image and multimodal generation, yet their deployment raises persistent questions about alignment, safety, preference satisfaction, and robustness to misuse. This survey reviews recent progress on aligning text-to-image diffusion models through reinforcement learning, reward modeling, preference optimization, and safety-specific fine-tuning. We organize the literature along five axes: the source of feedback, the form of the reward or preference signal, the optimization mechanism, the treatment of distribution shift and reward overoptimization, and the extent to which safety is addressed as an explicit constraint rather than a generic preference. The review covers reinforcement learning from human feedback, KL-regularized policy optimization, direct preference optimization, binary utility optimization, differentiable reward fine-tuning, surrogate reward learning, region-aware fine-tuning, and safety-oriented DPO variants. To make the survey accessible, we include tutorial explanations of diffusion sampling, reward modeling, and preference optimization, and briefly connect image diffusion alignment to emerging text and masked language diffusion models. We also compare representative methods in terms of feedback requirements, computational cost, scalability, susceptibility to reward hacking, and suitability for safety-critical deployment. Finally, we synthesize the literature into a set of open challenges: multi-objective alignment, feedback-efficient preference learning, adversarially robust safety alignment, continual alignment under changing norms, and interpretable reward modeling. The goal of this survey is to provide a coherent technical map of the emerging area of diffusion model alignment and to identify the methodological gaps that must be addressed before aligned generative models can be reliably deployed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。