用因果扩散模型从极低码率视频中重建高质量画面。
A Causal Diffusion Model for Video Reconstruction from Ultra-Low-Bitrate Representations
- 设计因果扩散模型融合语义与压缩帧信息联合重建。
- 在极低码率下优于传统、神经及生成类方法。
- 适合需要高效传输与高画质重建的场景。
我们研究从极低码率表示中重建视频,此时挑战从编码转向解码。传统和神经编码器重建会出现模糊,而生成与语义方法常难以同时保持保真度、时序一致性和感知质量。为此,我们提出一种因果视频扩散模型,通过联合建模互补的语义与高度压缩帧信息来重建视频。此外,引入仅时序的蒸馏机制,利用双向教师模型实现参数高效的训练与因果的少步推理。通过广泛的定量、定性和主观评估,证明该方法在极低码率视频重建中优于经典、神经、生成和语义基线。
原文摘要 · Abstract (English)
We study video reconstruction from ultra-low-bitrate representations, where the primary challenge shifts from encoding to decoding. In this regime, reconstruction with classical and neural codecs introduces blur, while generative and semantic approaches often struggle to jointly preserve fidelity, temporal consistency, and perceptual quality. To address these limitations, we propose a causal video diffusion model that reconstructs videos from ultra-low-bitrate semantics and highly compressed frames by jointly modeling their complementary information. We further introduce temporal-only distillation from a bidirectional teacher to enable parameter-efficient training and causal few-step inference. Through extensive quantitative, qualitative, and subjective evaluation, we show that the proposed method outperforms classical, neural, generative, and semantic baselines in ultra-low-bitrate video reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。