用参考图像增强视频生成细节和一致性
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding

- 在解码器中引入参考图像注意力,实现条件化视频重建
- 在多个基准上提升最高2.1dB PSNR,显著改善结构保真度
- 无需微调即可适配现有系统,适合视频生成与编辑场景
视频生成支撑众多下游应用。尽管主流潜空间扩散模型采用高度条件化的去噪网络,其解码器却通常保持无条件。我们观察到这种架构不对称会导致细节丢失与输入图像不一致。为此,我们提出RefDecoder——一种通过参考注意力直接注入高保真参考图像信号的条件化视频VAE解码器。轻量级图像编码器将参考帧映射为细节丰富的高维标记,在每个解码上采样阶段与去噪视频潜在标记共同处理。实验显示,该方法在Inter4K、WebVid和Large Motion重建基准上,对多种解码器骨干(如Wan 2.1和VideoVAE+)均实现稳定提升,最高达+2.1dB PSNR。RefDecoder可直接替换现有系统,无需额外微调,并在VBench I2V基准上全面提升主体一致性、背景一致性和整体质量评分。此外,该方法还广泛适用于风格迁移与视频编辑优化等任务。
原文摘要 · Abstract (English)
Video generation powers a vast array of downstream applications. However, while the de facto standard, i.e., latent diffusion models, typically employ heavily conditioned denoising networks, their decoders often remain unconditional. We observe that this architectural asymmetry leads to significant loss of detail and inconsistency relative to the input image. To address this, we argue that the decoder requires equal conditioning to preserve structural integrity. We introduce RefDecoder, a reference-conditioned video VAE decoder by injecting high-fidelity reference image signal directly into the decoding process via reference attention. Specifically, a lightweight image encoder maps the reference frame into the detail-rich high-dimensional tokens, which are co-processed with the denoised video latent tokens at each decoder up-sampling stage. We demonstrate consistent improvements across several distinct decoder backbones (e.g., Wan 2.1 and VideoVAE+), achieving up to +2.1dB PSNR over the unconditional baselines on the Inter4K, WebVid, and Large Motion reconstruction benchmarks. Notably, RefDecoder can be directly swapped into existing video generation systems without additional fine-tuning, and we report across-the-board improvements in subject consistency, background consistency, and overall quality scores on the VBench I2V benchmark. Beyond I2V, RefDecoder generalizes well to a wide range of visual generation tasks such as style transfer and video editing refinement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。