用对数编码让预训练模型直接生成高质量HDR视频。
HDR Video Generation via Latent Alignment with Logarithmic Encoding

- 用对数编码使HDR数据与预训练模型潜空间对齐,无需重训练编码器。
- 仅通过轻量微调即可生成高动态范围视频,跨场景表现稳定。
- 适合想快速实现HDR生成的开发者和影视制作人员。
高动态范围(HDR)图像能更真实地还原场景辐射度,但因其与生成模型训练所用有界、感知压缩数据不匹配,难以直接应用。本文提出一种简单方法:利用预训练生成模型已学习的强视觉先验。我们发现电影制作中常用的对数编码可将HDR图像映射到与模型潜空间自然对齐的分布,从而通过轻量微调实现直接适配,无需重新训练编码器。为进一步恢复输入中不可见的细节,我们设计了一种基于相机模拟退化的训练策略,促使模型从其先验中推断缺失的高动态范围内容。结合上述思路,仅以少量调整即可在预训练视频模型上实现高质量HDR视频生成,在多样场景和复杂光照条件下均表现优异。结果表明,只要选择合适的表示方式,即使在根本不同的成像机制下,也无需重构生成模型即可有效处理HDR内容。
原文摘要 · Abstract (English)
High dynamic range (HDR) imagery offers a rich and faithful representation of scene radiance, but remains challenging for generative models due to its mismatch with the bounded, perceptually compressed data on which these models are trained. A natural solution is to learn new representations for HDR, which introduces additional complexity and data requirements. In this work, we show that HDR generation can be achieved in a much simpler way by leveraging the strong visual priors already captured by pretrained generative models. We observe that a logarithmic encoding widely used in cinematic pipelines maps HDR imagery into a distribution that is naturally aligned with the latent space of these models, enabling direct adaptation via lightweight fine-tuning without retraining an encoder. To recover details that are not directly observable in the input, we further introduce a training strategy based on camera-mimicking degradations that encourages the model to infer missing high dynamic range content from its learned priors. Combining these insights, we demonstrate high-quality HDR video generation using a pretrained video model with minimal adaptation, achieving strong results across diverse scenes and challenging lighting conditions. Our results indicate that HDR, despite representing a fundamentally different image formation regime, can be handled effectively without redesigning generative models, provided that the representation is chosen to align with their learned priors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。