不重训练,通过调节采样过程实现无损高动态范围图像生成
LumaGuide: Distribution Shaping for Training-Free HDR Generation in Diffusion Models

- 在采样阶段用可微能量引导控制亮度分布,无需修改模型参数
- 对亮度直方图进行匹配,即可生成具有清晰高光与暗部细节的HDR图像
- 支持文本、参考图或预设数据驱动的灵活分布控制,适合视频生成
预训练扩散模型生成图像逼真,但受限于训练数据的统计偏差,难以生成高动态范围(HDR)内容。本文提出LumaGuide,一种无需训练的扩散模型分布调节框架。该框架不修改模型参数,而是通过可微的能量引导机制,在采样过程中调节输出特征分布以匹配目标分布。针对HDR生成,我们基于感知均匀的PQ空间控制亮度分布。实验表明,仅对亮度直方图进行对齐,即可诱导出一致的高光表现和保留阴影细节,同时保持语义一致性。该方法还可通过数据驱动预设、参考图像或文本驱动预测器,灵活指定目标分布,并自然扩展至具备时间一致性的视频生成。更广泛地,本工作证明:可通过采样阶段直接塑造输出分布,实现可控生成而无需重新训练扩散模型。
原文摘要 · Abstract (English)
Pretrained diffusion models generate realistic images but are constrained by the statistical biases of their training data, limiting their ability to produce high dynamic range (HDR) content. In this work, we introduce LumaGuide, a training-free framework for distribution shaping in diffusion models. Instead of modifying model parameters, LumaGuide steers the sampling process to match target feature distributions via differentiable energy-based guidance. We instantiate this framework for HDR generation by controlling luminance distributions in perceptually uniform PQ space. Our results show that aligning luminance histograms is sufficient to induce HDR-consistent behavior, including coherent highlights and preserved shadow detail, while maintaining semantic fidelity. Beyond HDR, LumaGuide enables flexible specification of target distributions through data-driven presets, reference images, or text-driven predictors, and extends naturally to video generation with temporal consistency constraints. More broadly, our work demonstrates that controllable generation can be achieved by directly shaping output distributions at sampling time, without retraining diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。