用视频扩散模型统一压缩与修复,解决生成式视频编码的闪烁问题。
Generative Neural Video Compression via Video Diffusion Prior
- 基于视频扩散变换器,统一时空潜在表示压缩与序列级去噪修复。
- 在低于0.01 bpp时仍显著减少闪烁,感知质量超越传统与学习型编码器。
- 适合追求极致低码率下画质稳定的视频压缩研究者与工程师。
我们提出GNVC-VD,首个基于DiT的生成式神经视频压缩框架,依托先进的视频生成基础模型,将时空潜在压缩与序列级生成优化统一于单一编解码器中。现有感知编码器主要依赖预训练图像生成先验恢复高频细节,但其逐帧处理缺乏时序建模,导致感知闪烁。为解决此问题,GNVC-VD引入统一的流匹配潜在修复模块,利用视频扩散变换器通过序列级去噪联合增强帧内与帧间潜在表示,确保时空细节一致性。不同于视频生成中从纯高斯噪声开始去噪,GNVC-VD从解码后的时空潜在表示初始化,并学习一个适应压缩退化特征的修正项,使扩散先验适配压缩损伤。此外,条件适配器向DiT中间层注入压缩感知线索,有效去除伪影同时保持极端低码率下的时序连贯性。大量实验表明,GNVC-VD在感知质量上超越传统及学习型编码器,显著降低先前生成方法中持续存在的闪烁现象,即使在低于0.01 bpp时亦然,凸显了将视频原生生成先验融入神经编解码器在下一代感知视频压缩中的潜力。
原文摘要 · Abstract (English)
We present GNVC-VD, the first DiT-based generative neural video compression framework built upon an advanced video generation foundation model, where spatio-temporal latent compression and sequence-level generative refinement are unified within a single codec. Existing perceptual codecs primarily rely on pre-trained image generative priors to restore high-frequency details, but their frame-wise nature lacks temporal modeling and inevitably leads to perceptual flickering. To address this, GNVC-VD introduces a unified flow-matching latent refinement module that leverages a video diffusion transformer to jointly enhance intra- and inter-frame latents through sequence-level denoising, ensuring consistent spatio-temporal details. Instead of denoising from pure Gaussian noise as in video generation, GNVC-VD initializes refinement from decoded spatio-temporal latents and learns a correction term that adapts the diffusion prior to compression-induced degradation. A conditioning adaptor further injects compression-aware cues into intermediate DiT layers, enabling effective artifact removal while maintaining temporal coherence under extreme bitrate constraints. Extensive experiments show that GNVC-VD surpasses both traditional and learned codecs in perceptual quality and significantly reduces the flickering artifacts that persist in prior generative approaches, even below 0.01 bpp, highlighting the promise of integrating video-native generative priors into neural codecs for next-generation perceptual video compression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。