用视频生成思路重建高动态范围图像,更准且可解释。
Single-Shot HDR Recovery via a Video Diffusion Prior

- 将单张图像HDR重建转为条件视频生成,先生成多曝光帧。
- 融合时用轻量UNet预测每像素权重,保持输入真实性。
- 方法通用,可扩展至模糊图像聚焦等任务,适合图像重建研究者。
近期基于生成模型的单张高动态范围(HDR)图像重建方法表现良好,但常难以保持与输入图像的一致性。这些方法通常需分别处理亮部和暗部,或直接预测最终HDR图像而牺牲可解释性。本文将单张HDR重建重新建模为条件视频生成任务,通过微调视频扩散模型生成以低动态范围(LDR)输入为条件的多曝光图像序列,并利用轻量级UNet预测每像素权重进行融合。该方法结构简单、可解释性强且效果优异。不同于直接‘幻想’出HDR图像,本方法显式重建中间曝光堆栈并融合输出。无需针对不同曝光区间使用独立模型,且重建结果具有更高输入保真度。在多个定量基准上,本方法在与现有生成基线相当的模型容量下,各项重建指标均优于当前最优。人工评估显示,在72%的成对比较中,人类更偏好本方法的结果。此外,该输入条件化序列生成与融合框架还可推广至其他图像重建任务,例如从单张散焦模糊图像恢复全清晰图像。
原文摘要 · Abstract (English)
Recent generative methods for single-shot high dynamic range (HDR) image reconstruction show promising results, but often struggle with preserving fidelity to the input image. They require separate models to handle highlights and shadows, or sacrifice interpretability by directly predicting the final HDR image. We address these limitations by re-casting single-shot HDR reconstruction as conditional video generation and fusing the generated frames into an HDR image. We finetune a video diffusion model to generate an exposure bracket, conditioned on a low dynamic range (LDR) input. We fuse this image bracket using per-pixel weights predicted by a light-weight UNet. This formulation is simple, interpretable, and effective. Rather than directly hallucinating an HDR image, it explicitly reconstructs the intermediate exposure stack and fuses it into the final output. Our method eliminates the need for separate models across exposure regimes and produces HDR reconstructions with high input fidelity. On quantitative benchmarks, we outperform state-of-the-art generative baselines with comparable model capacity on several reconstruction metrics. Human evaluators further prefer our results in 72% of pairwise comparisons against existing methods. Finally, we show that this input-conditioned sequence generation and fusion framework extends beyond HDR to other image reconstruction tasks, such as all-in-focus image recovery from a single defocus-blurred input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。