arXiv:2604.14379cs.LGcs.AI2026-04

无需重训练,一步到位实现多目标图像生成对齐。

Step-level Denoising-time Diffusion Alignment with Multiple Objectives

论文配图:Step-level Denoising-time Diffusion Alignment with Multiple Objectives
图 1 · 摘自论文原文
  • 按步骤设计强化学习框架,直接求解最优去噪分布。
  • 理论证明去噪目标与强化学习完全等价,无近似误差。
  • 适合需要兼顾美感与图文一致性的生成模型优化场景。

强化学习(RL)已成为对齐扩散模型与人类偏好的有力工具,通常在KL正则化约束下优化单一奖励函数。然而,人类偏好本质上是多元的,对齐模型需平衡多个下游目标,如美学质量与文本-图像一致性。现有方法要么依赖昂贵的多目标RL微调,要么在去噪时融合独立对齐的模型,但通常需要奖励值(或其梯度),或引入去噪目标的近似误差。本文重新审视扩散模型的强化学习微调问题,通过引入步骤级(step-level)RL公式,解决了最优策略难以识别的难题。在此基础上,提出多目标步骤级去噪时间扩散对齐(MSDDA)框架,无需重训练即可实现多目标对齐,以闭式表达获得最优反向去噪分布,其均值与方差可直接由单目标基础模型表示。我们证明该去噪时间目标与步骤级强化学习微调完全等价,不引入近似误差。数值结果表明,该方法优于现有去噪时间方法。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has emerged as a powerful tool for aligning diffusion models with human preferences, typically by optimizing a single reward function under a KL regularization constraint. In practice, however, human preferences are inherently pluralistic, and aligned models must balance multiple downstream objectives, such as aesthetic quality and text-image consistency. Existing multi-objective approaches either rely on costly multi-objective RL fine-tuning or on fusing separately aligned models at denoising time, but they generally require access to reward values (or their gradients) and/or introduce approximation error in the resulting denoising objectives. In this paper, we revisit the problem of RL fine-tuning for diffusion models and address the intractability of identifying the optimal policy by introducing a step-level RL formulation. Building on this, we further propose Multi-objective Step-level Denoising-time Diffusion Alignment (MSDDA), a retraining-free framework for aligning diffusion models with multiple objectives, obtaining the optimal reverse denoising distribution in closed form, with mean and variance expressed directly in terms of single-objective base models. We prove that this denoising-time objective is exactly equivalent to the step-level RL fine-tuning, introducing no approximation error. Moreover, we provide numerical results, which indicate our method outperforms existing denoising-time approaches.

扩散模型多目标对齐强化学习去噪时间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。