用单步扩散模型实现低码率高效视频编码,提升画质同时降低计算开销。
Single-step Diffusion-based Video Coding with Semantic-Temporal Guidance
- 采用条件编码与单步扩散生成器结合,实现快速高质量重建。
- 在低码率下相比现有感知方法平均节省52.73%码率,保持视觉真实感。
- 引入语义-时序引导机制,增强帧间一致性与生成细节表现。
尽管传统和神经视频编码器(NVCs)在率失真性能上已取得显著进展,但在低码率下提升主观质量仍具挑战。部分NVCs虽引入感知或对抗目标,但受限于生成能力仍存在伪影;另一些则利用预训练扩散模型提升质量,却带来高昂采样复杂度。为此,我们提出S2VC——一种基于单步扩散的视频编码器,融合条件编码框架与高效单步扩散生成器,可在低码率下实现逼真重建并降低采样成本。针对单步扩散中语义条件的重要性,我们提出上下文语义引导,从缓冲特征中提取帧自适应语义,替代文本提示,提升生成真实性。此外,在扩散U-Net中引入时序一致性引导,强化帧间连续性,确保生成稳定性。大量实验表明,S2VC在感知质量上达到当前最优,相比先前感知方法平均节省52.73%码率,凸显单步扩散在高效高质视频压缩中的潜力。
原文摘要 · Abstract (English)
While traditional and neural video codecs (NVCs) have achieved remarkable rate-distortion performance, improving perceptual quality at low bitrates remains challenging. Some NVCs incorporate perceptual or adversarial objectives but still suffer from artifacts due to limited generation capacity, whereas others leverage pretrained diffusion models to improve quality at the cost of heavy sampling complexity. To overcome these challenges, we propose S2VC, a Single-Step diffusion based Video Codec that integrates a conditional coding framework with an efficient single-step diffusion generator, enabling realistic reconstruction at low bitrates with reduced sampling cost. Recognizing the importance of semantic conditioning in single-step diffusion, we introduce Contextual Semantic Guidance to extract frame-adaptive semantics from buffered features. It replaces text captions with efficient, fine-grained conditioning, thereby improving generation realism. In addition, Temporal Consistency Guidance is incorporated into the diffusion U-Net to enforce temporal coherence across frames and ensure stable generation. Extensive experiments show that S2VC delivers state-of-the-art perceptual quality with an average 52.73% bitrate saving over prior perceptual methods, underscoring the promise of single-step diffusion for efficient, high-quality video compression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。