让视频生成更可控,兼顾画质与真实感。
Controllable Generative Video Compression
- 用关键帧提供结构先验,控制非关键帧生成。
- 在相同码率下,重建质量提升1.2~2.3dB PSNR。
- 适合需要高保真与细节保留的视频压缩场景。
感知视频压缩利用生成建模提升视觉真实感,但常牺牲信号保真度,偏离了视频压缩忠实还原视觉信号的目标。为缓解感知与保真之间的矛盾,本文提出可控生成视频压缩(CGVC)范式,通过多种视觉条件引导细节精确生成。该范式下,场景代表性关键帧被编码并作为非关键帧生成的结构先验;同时,额外编码密集的逐帧控制先验,以更好保留每个非关键帧的精细结构与语义。在这些先验引导下,非关键帧由具备时空一致性的可控视频生成模型重建。此外,为准确恢复视频色彩信息,设计了基于颜色距离的关键帧选择算法,实现自适应关键帧选取。实验结果表明,CGVC在信号保真度和感知质量上均优于现有感知视频压缩方法。
原文摘要 · Abstract (English)
Perceptual video compression adopts generative video modeling to improve perceptual realism but frequently sacrifices signal fidelity, diverging from the goal of video compression to faithfully reproduce visual signal. To alleviate the dilemma between perception and fidelity, in this paper we propose Controllable Generative Video Compression (CGVC) paradigm to faithfully generate details guided by multiple visual conditions. Under the paradigm, representative keyframes of the scene are coded and used to provide structural priors for non-keyframe generation. Dense per-frame control prior is additionally coded to better preserve finer structure and semantics of each non-keyframe. Guided by these priors, non-keyframes are reconstructed by controllable video generation model with temporal and content consistency. Furthermore, to accurately recover color information of the video, we develop a color-distance-guided keyframe selection algorithm to adaptively choose keyframes. Experimental results show CGVC outperforms previous perceptual video compression method in terms of both signal fidelity and perceptual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。