用场景动态先验压缩视频,低码率下仍能高清还原。
Tokenizing Motion: A Generative Approach for Scene Dynamics Compression
- 用自然场景的细微运动模式做压缩先验,替代传统内容先验。
- 在极低码率下实现高质量动态重建,优于现有编码器。
- 适合需要高效传输复杂场景视频的实时应用。
本文提出一种新型生成式视频压缩框架,利用常见场景中细微动态(如摇曳的花朵或漂浮的船只)所蕴含的运动模式先验,而非依赖传统视频内容先验(如说话人脸或人体动作)。这些紧凑的运动先验使超低码率通信成为可能,并在多种场景内容下实现高质量重建。编码端通过密集到稀疏的变换将运动先验压缩为紧凑表示;解码端则利用先进的流驱动扩散模型重建场景动态。实验表明,该方法在速率-失真性能上优于当前最先进的传统视频编码器Enhanced Compression Model (ECM),尤其在场景动态序列上表现突出。项目页面见:https://github.com/xyzysz/GNVDC。
原文摘要 · Abstract (English)
This paper proposes a novel generative video compression framework that leverages motion pattern priors, derived from subtle dynamics in common scenes (e.g., swaying flowers or a boat drifting on water), rather than relying on video content priors (e.g., talking faces or human bodies). These compact motion priors enable a new approach to ultra-low bitrate communication while achieving high-quality reconstruction across diverse scene contents. At the encoder side, motion priors can be streamlined into compact representations via a dense-to-sparse transformation. At the decoder side, these priors facilitate the reconstruction of scene dynamics using an advanced flow-driven diffusion model. Experimental results illustrate that the proposed method can achieve superior rate-distortion-performance and outperform the state-of-the-art conventional-video codec Enhanced Compression Model (ECM) on-scene dynamics sequences. The project page can be found at-https://github.com/xyzysz/GNVDC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。