arXiv:2507.15269cs.CVcs.AI2025-07被引 5

用条件扩散模型生成视频,提升高压缩下的视觉质量。

Conditional Video Generation for High-Efficiency Video Compression

  • 将视频压缩转为条件生成任务,从稀疏信号重建视频
  • 在高压缩比下FVD和LPIPS指标显著优于传统与神经编码器
  • 适合关注感知质量与高效传输的视频压缩研究者

感知研究表明,条件扩散模型在符合人类视觉感知的视频重建方面表现优异。基于此,我们提出一种利用条件扩散模型实现感知优化重建的视频压缩框架。具体而言,将视频压缩重构为条件生成任务,由生成模型从稀疏但信息丰富的信号中合成视频。本方法引入三个关键模块:(1) 多粒度条件建模,捕捉静态场景结构与动态时空特征;(2) 高效传输的紧凑表示,不损失语义丰富性;(3) 多条件训练结合模态丢弃与角色感知嵌入,避免对单一模态依赖,提升鲁棒性。大量实验表明,该方法在高压缩比下显著优于传统与神经编码器,在感知质量指标如弗雷谢视频距离(FVD)和LPIPS上表现突出。

原文摘要 · Abstract (English)

Perceptual studies demonstrate that conditional diffusion models excel at reconstructing video content aligned with human visual perception. Building on this insight, we propose a video compression framework that leverages conditional diffusion models for perceptually optimized reconstruction. Specifically, we reframe video compression as a conditional generation task, where a generative model synthesizes video from sparse, yet informative signals. Our approach introduces three key modules: (1) Multi-granular conditioning that captures both static scene structure and dynamic spatio-temporal cues; (2) Compact representations designed for efficient transmission without sacrificing semantic richness; (3) Multi-condition training with modality dropout and role-aware embeddings, which prevent over-reliance on any single modality and enhance robustness. Extensive experiments show that our method significantly outperforms both traditional and neural codecs on perceptual quality metrics such as Fréchet Video Distance (FVD) and LPIPS, especially under high compression ratios.

视频生成扩散模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。