arXiv:2608.20515cs.CV2026-08

用单步扩散模型实现低延迟高保真视频压缩

DiffVC-ONE: Diffusion-based Generative Video Compression with One-Step Video Diffusion Transformer

论文配图:DiffVC-ONE: Diffusion-based Generative Video Compression with One-Step Video Diffusion Transformer
图 1 · 摘自论文原文
  • 单步视频扩散变换器+统一单向潜码流压缩
  • 比现有方法提升视觉质量,推理成本更低
  • 适合实时视频传输与高质量编码场景

生成式视频压缩可在低码率下恢复丰富视觉细节,但同时实现高时序一致性与低推理开销仍具挑战。为此,我们提出 DiffVC-ONE,一种基于单步视频扩散变换器的生成式视频压缩框架。首先,引入统一单向潜码流压缩器,使用共享模型高效统一地压缩紧凑潜码流。随后,设计基于视频扩散变换器的单步增强模块,以重建潜码流为内容锚点,对整个图像组进行单步时空感知增强。最后,混合条件生成器从重建内容与量化信息中提取结构、强度和语义条件,以保留真实区域、控制生成增强程度,并在单步扩散增强中补充内容感知的感知细节。在多个标准基准上的大量实验表明,DiffVC-ONE 在保持低推理成本的同时,实现了最先进的感知质量与时序一致性。

原文摘要 · Abstract (English)

Generative video compression can recover rich visual details at low bitrates, but simultaneously achieving high temporal consistency and low inference cost remains challenging. To address this issue, we propose DiffVC-ONE, a diffusion-based generative video compression framework built on a one-step Video Diffusion Transformer. First, we introduce a Unified Unidirectional Latent Compressor that uses a shared model to efficiently and uniformly compress compact latent slices. We then develop a Video DiT-based One-Step Diffusion Enhancer that uses the reconstructed latent slices as content anchors and performs single-step spatio-temporal perceptual enhancement over an entire group of pictures. Finally, a Hybrid Condition Generator extracts structural, strength, and semantic conditions from the reconstructed content and quantization information. These conditions preserve faithful regions, control the degree of generative enhancement, and supplement content-aware perceptual details during one-step diffusion enhancement. Extensive experiments on multiple standard benchmarks demonstrate that DiffVC-ONE achieves state-of-the-art perceptual quality and temporal consistency with low inference cost.

视频压缩扩散模型生成式编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。