无需训练的视频压缩框架,显著提升超低码率下的时序连贯性。
Free-GVC: Towards Training-Free Extreme Generative Video Compression with Temporal Coherence
- 基于扩散模型先验,将视频编码转为潜在轨迹压缩。
- 在超低码率下比最新神经编码器DCVC-RT降低93.29%的BD-Rate。
- 适合追求高感知质量与时序稳定性的视频压缩应用。
基于视频生成的最新进展,生成式视频压缩成为实现视觉愉悦重建的新范式。然而,现有方法对时序相关性的利用有限,在超低码率下导致明显闪烁和时序连贯性下降。本文提出Free-GVC,一种无需训练的生成式视频压缩框架,将视频编码重新定义为由视频扩散先验引导的潜在轨迹压缩。该方法在图像组(GOP)层面操作,将视频片段编码至紧凑潜在空间,并沿扩散轨迹逐步压缩。为确保跨GOP的感知一致性,引入自适应质量控制模块,动态构建在线码率-感知代理模型,预测每组的最佳扩散步数。此外,跨GOP对齐模块建立帧重叠并执行潜在融合,有效缓解闪烁,增强时序连贯性。实验表明,Free-GVC在DISTS指标上相比最新神经编码器DCVC-RT平均降低93.29%的BD-Rate,用户研究进一步证实其在超低码率下具备更优的感知质量和时序连贯性。
原文摘要 · Abstract (English)
Building on recent advances in video generation, generative video compression has emerged as a new paradigm for achieving visually pleasing reconstructions. However, existing methods exhibit limited exploitation of temporal correlations, causing noticeable flicker and degraded temporal coherence at ultra-low bitrates. In this paper, we propose Free-GVC, a training-free generative video compression framework that reformulates video coding as latent trajectory compression guided by a video diffusion prior. Our method operates at the group-of-pictures (GOP) level, encoding video segments into a compact latent space and progressively compressing them along the diffusion trajectory. To ensure perceptually consistent reconstruction across GOPs, we introduce an Adaptive Quality Control module that dynamically constructs an online rate-perception surrogate model to predict the optimal diffusion step for each GOP. In addition, an Inter-GOP Alignment module establishes frame overlap and performs latent fusion between adjacent groups, thereby mitigating flicker and enhancing temporal coherence. Experiments show that Free-GVC achieves an average of 93.29% BD-Rate reduction in DISTS over the latest neural codec DCVC-RT, and a user study further confirms its superior perceptual quality and temporal coherence at ultra-low bitrates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。