arXiv:2512.00408cs.CVcs.AI2025-12被引 3

用语义引导的生成模型,实现超低码率下的高清视频压缩

Low-Bitrate Video Compression through Semantic-Conditioned Diffusion

  • 将视频分解为语义、外观和运动三类紧凑信息,仅传输关键内容
  • 在极低码率下,感知质量提升2至10倍,显著优于传统编码方法
  • 适合需要极致压缩比的视频传输场景,如移动流媒体或远程监控

传统视频编码器在超低码率下因追求像素精度而产生严重伪影,根源在于像素准确与人眼感知不匹配。本文提出名为DiSCo的语义视频压缩框架,仅传输最具意义的信息,并依赖生成先验恢复细节。原始视频被分解为三类紧凑模态:文本描述、时空退化的视频,以及可选的草图或姿态信息,分别捕捉语义、外观和运动线索。通过条件视频扩散模型,从这些紧凑表示中重建高质量、时间连贯的视频。设计了时间前向填充、标记交错和模态专用编码器,以提升多模态生成效果与模态紧凑性。实验表明,在低码率下,本方法在感知指标上优于基线语义编码和传统编码2至10倍。

原文摘要 · Abstract (English)

Traditional video codecs optimized for pixel fidelity collapse at ultra-low bitrates and produce severe artifacts. This failure arises from a fundamental misalignment between pixel accuracy and human perception. We propose a semantic video compression framework named DiSCo that transmits only the most meaningful information while relying on generative priors for detail synthesis. The source video is decomposed into three compact modalities: a textual description, a spatiotemporally degraded video, and optional sketches or poses that respectively capture semantic, appearance, and motion cues. A conditional video diffusion model then reconstructs high-quality, temporally coherent videos from these compact representations. Temporal forward filling, token interleaving, and modality-specific codecs are proposed to improve multimodal generation and modality compactness. Experiments show that our method outperforms baseline semantic and traditional codecs by 2-10X on perceptual metrics at low bitrates.

视频压缩扩散模型语义编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。