arXiv:2604.23508cs.CV2026-04

用视频生成先验提升多帧低分辨率图像的超分辨率质量

BurstGP: Enhancing Raw Burst Image Super Resolution with Generative Priors

论文配图:BurstGP: Enhancing Raw Burst Image Super Resolution with Generative Priors
图 1 · 摘自论文原文
  • 基于扩散模型融合视频级生成先验,增强多帧图像重建
  • 在MUSIQ和LPIPS指标上超越现有方法,纹理更丰富
  • 适合追求真实细节与视觉质量的图像超分研究者

多帧图像超分辨率(BISR)旨在通过聚合多个低分辨率帧的信息,重建单个高分辨率图像,依赖于序列中的时间冗余与空间一致性。传统方法虽表现良好,但在复杂纹理和细节恢复上常出现过平滑问题。扩散模型,尤其是预训练于高质量数据的模型,在图像与视频超分辨率中展现出生成逼真细节的强大能力,但其在BISR中的应用仍不充分。现有方法多采用从零训练的任务特定扩散模型,且仅处理单帧重建。本文提出BurstGP,一种基于扩散模型的BISR新方法,利用近期基础模型的生成先验来克服上述挑战。具体而言,我们在传统BISR框架之上构建了多帧感知扩散模型,以极小损失保持保真度的同时显著提升图像质量。此外,提出:(i) 一种新的退化感知条件机制,根据输入估计的退化程度控制细节合成;(ii) 一个鲁棒的sRGB-to-lRGB逆变换器,使我们能使用视频级sRGB生成先验,同时处理原始输入和lRGB输出。实验表明,BurstGP在定量(尤其在感知指标如MUSIQ和LPIPS上)和定性结果上均优于现有最佳方法,尤其在恢复更丰富的纹理和细微结构方面表现突出,凸显视频先验在BISR中的潜力。

原文摘要 · Abstract (English)

Burst image super resolution (BISR) aims to construct a single high-resolution (HR) image by aggregating information from multiple low-resolution (LR) frames, relying on temporal redundancy and spatial coherence across the burst. While conventional methods achieve impressive results, they often struggle with complex textures and oversmoothing. Diffusion models, particularly those pretrained on high-quality data, have shown remarkable capability in generating realistic details for image and video super-resolution. However, their potential remains largely under-explored in BISR, where existing approaches typically rely on task-specific diffusion models trained from scratch and operate on single-frame reconstructions. In this work, we propose BurstGP, a novel diffusion-based solution for BISR, which leverages generative priors of recent foundation models to overcome these issues. In particular, we build a multiframe-aware diffusion model on top of a conventional BISR approach, which boosts image quality with minimal loss to fidelity. Further, we introduce (i) a novel degradation-aware conditioning mechanism, which controls synthesis of fine details based on the estimated degradation in the input, and (ii) a robust sRGB-to-lRGB inverter, enabling us to utilize generative multiframe (video) sRGB priors, while operating with raw input and lRGB output images. Empirically, we demonstrate that BurstGP outperforms the existing state of the art, both quantitatively (especially with respect to perceptual metrics, including MUSIQ and LPIPS) and qualitatively. In particular, our proposed method excels at recovering richer textures and finer structural details, highlighting the potential of video priors for BISR over traditional methods.

图像超分扩散模型视频先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。