arXiv:2602.05201eess.IV2026-02中稿 · ICASSP 2026被引 2

用扩散模型压缩视频,在极低码率下保持画质

Diffusion-aided Extreme Video Compression with Lightweight Semantics Guidance

  • 用语义特征代替原始像素,通过条件扩散模型重建画面
  • 仅需500 bit/帧即可实现高质量视频重建
  • 适合对带宽敏感的实时视频传输场景

现代视频编码器和基于学习的方法在极低码率下难以实现语义重建,因其依赖低层次时空冗余。生成模型,尤其是扩散模型,通过利用高层语义理解与强大图像合成能力,为视频压缩提供了新范式。本文提出一种视频压缩框架,整合生成先验以大幅降低码率,同时保持重建保真度。具体而言,该方法压缩视频的高层语义表示,再使用条件扩散模型从这些语义中重构帧。为进一步提升压缩效率,采用全局相机轨迹和前景分割表征运动信息:背景运动由相机姿态参数紧凑表示,前景动态则由稀疏分割掩码描述。该设计显著提升压缩效率,使极低码率下的高质量视频重建成为可能。

原文摘要 · Abstract (English)

Modern video codecs and learning-based approaches struggle for semantic reconstruction at extremely low bit-rates due to reliance on low-level spatiotemporal redundancies. Generative models, especially diffusion models, offer a new paradigm for video compression by leveraging high-level semantic understanding and powerful visual synthesis. This paper propose a video compression framework that integrates generative priors to drastically reduce bit-rate while maintaining reconstruction fidelity. Specifically, our method compresses high-level semantic representations of the video, then uses a conditional diffusion model to reconstruct frames from these semantics. To further improve compression, we characterize motion information with global camera trajectories and foreground segmentation: background motion is compactly represented by camera pose parameters while foreground dynamics by sparse segmentation masks. This allows for significantly boosts compression efficiency, enabling descent video reconstruction at extremely low bit-rates.

视频压缩扩散模型语义编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。