用单步扩散模型实现更快更清晰的视频压缩,感知质量显著提升。
DiffVC-OSD: One-Step Diffusion-based Perceptual Neural Video Compression Framework
- 单步扩散直接重构潜在表示,结合时序上下文提升画质。
- 相比多步方法,解码速度提升20倍,码率降低86.92%。
- 适合追求高速高质视频压缩的研究与应用开发者。
本文提出 DiffVC-OSD,一种基于单步扩散的感知神经视频压缩框架。不同于传统多步扩散方法,DiffVC-OSD 将重建后的潜在表示直接输入单步扩散模型,通过时序上下文和潜在特征共同引导,提升感知质量。为更好利用时序依赖,设计了时序上下文适配器,将条件输入编码为多层次特征,为去噪 U-Net 提供更精细的引导。同时采用端到端微调策略优化整体压缩性能。大量实验表明,DiffVC-OSD 在感知压缩性能上达到当前最优,相比对应多步扩散基线,解码速度提升约20倍,码率降低86.92%。
原文摘要 · Abstract (English)
In this work, we first propose DiffVC-OSD, a One-Step Diffusion-based Perceptual Neural Video Compression framework. Unlike conventional multi-step diffusion-based methods, DiffVC-OSD feeds the reconstructed latent representation directly into a One-Step Diffusion Model, enhancing perceptual quality through a single diffusion step guided by both temporal context and the latent itself. To better leverage temporal dependencies, we design a Temporal Context Adapter that encodes conditional inputs into multi-level features, offering more fine-grained guidance for the Denoising Unet. Additionally, we employ an End-to-End Finetuning strategy to improve overall compression performance. Extensive experiments demonstrate that DiffVC-OSD achieves state-of-the-art perceptual compression performance, offers about 20$\times$ faster decoding and a 86.92\% bitrate reduction compared to the corresponding multi-step diffusion-based variant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。