arXiv:2506.12036cs.LGcs.AI2025-06被引 2

通过优化初始噪声提升文生图模型质量,无需修改主模型

A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models

  • 仅训练初始噪声生成器,保持预训练模型完全冻结
  • 在低推理步数下显著提升图文对齐与图像质量
  • 方法极简,无需轨迹存储或复杂奖励机制,适合高效微调

近期研究采用强化学习(RL)微调文生图扩散模型,改善图文对齐与生成质量。但现有方法引入过多复杂性:需缓存完整采样轨迹、依赖可微分奖励模型或大规模偏好数据集,或需要特殊引导技巧。受‘黄金噪声’假设启发——某些初始噪声样本能持续产生更优对齐结果——我们提出 Noise PPO,一种极简的强化学习算法,仅冻结预训练扩散模型,学习一个提示条件的初始噪声生成器。该方法无需轨迹存储、奖励反向传播或复杂引导技术。大量实验表明,优化初始噪声分布能持续提升模型对齐度与样本质量,尤其在低推理步数时收益最大。随着推理步数增加,优化收益递减但仍存在。这些结果明确了黄金噪声假设的适用范围与局限,证实了极简强化学习微调在扩散模型中的实用价值。

原文摘要 · Abstract (English)

Recent work uses reinforcement learning (RL) to fine-tune text-to-image diffusion models, improving text-image alignment and sample quality. However, existing approaches introduce unnecessary complexity: they cache the full sampling trajectory, depend on differentiable reward models or large preference datasets, or require specialized guidance techniques. Motivated by the "golden noise" hypothesis -- that certain initial noise samples can consistently yield superior alignment -- we introduce Noise PPO, a minimalist RL algorithm that leaves the pre-trained diffusion model entirely frozen and learns a prompt-conditioned initial noise generator. Our approach requires no trajectory storage, reward backpropagation, or complex guidance tricks. Extensive experiments show that optimizing the initial noise distribution consistently improves alignment and sample quality over the original model, with the most significant gains at low inference steps. As the number of inference steps increases, the benefit of noise optimization diminishes but remains present. These findings clarify the scope and limitations of the golden noise hypothesis and reinforce the practical value of minimalist RL fine-tuning for diffusion models.

文生图扩散模型强化学习微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。