用神经隐式表示+扩散模型,在极低码率下实现高质量视频压缩。
DiV-INR: Extreme Low-Bitrate Diffusion Video Compression with INR Conditioning
- 用INR替代关键帧,以神经编码方式高效引导扩散模型生成视频。
- 在0.05 bpp以下码率下,LPIPS、FID等感知指标显著优于HEVC和VVC。
- 适合追求极致压缩率且注重视觉感知质量的视频编码研究者。
我们提出一种感知驱动的视频压缩框架,结合隐式神经表征(INRs)与预训练视频扩散模型,应对极端低码率场景(<0.05 bpp)。该方法利用INRs的紧凑表征能力与扩散模型从大规模数据中学习到的丰富生成先验。通过用INR条件化取代传统内编码关键帧,仅以少量参数即可估计潜在特征并引导扩散过程。对INR权重与扩散模型的参数高效适配器进行联合优化,使模型在极小参数开销下学习可靠条件信号,并编码视频特异性信息。在UVG、MCL-JCV和JVET Class-B基准测试中,该方法在极低码率下显著提升感知质量(LPIPS、DISTS、FID),BD-LPIPS最高提升0.214,BD-FID最高提升91.14,超越HEVC、VVC及现有先进神经与INR-only视频编解码器。分析表明,该方法先生成场景布局与物体身份,再细化纹理细节,揭示了从语义到视觉的层级生成机制,从而实现极低码率下的感知保真压缩。
原文摘要 · Abstract (English)
We present a perceptually-driven video compression framework integrating implicit neural representations (INRs) and pre-trained video diffusion models to address the extremely low bitrate regime (<0.05 bpp). Our approach exploits the complementary strengths of INRs, which provide a compact video representation, and diffusion models, which offer rich generative priors learned from large-scale datasets. The INR-based conditioning replaces traditional intra-coded keyframes with bit-efficient neural representations trained to estimate latent features and guide the diffusion process. Our joint optimization of INR weights and parameter-efficient adapters for diffusion models allows the model to learn reliable conditioning signals while encoding video-specific information with minimal parameter overhead. Our experiments on UVG, MCL-JCV, and JVET Class-B benchmarks demonstrate substantial improvements in perceptual metrics (LPIPS, DISTS, and FID) at extremely low bitrates, including improvements on BD-LPIPS up to 0.214 and BD-FID up to 91.14 relative to HEVC, while also outperforming VVC and previous strong state-of-the-art neural and INR-only video codecs. Moreover, our analysis shows that INR-conditioned diffusion-based video compression first composes the scene layout and object identities before refining textural accuracy, exposing the semantic-to-visual hierarchy that enables perceptually faithful compression at extremely low bitrates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。