arXiv:2411.15798eess.IVcs.CV2024-11中稿 · ICASSP 2025被引 10

用多模态生成模型实现超低码率下可控视频压缩,保真度显著提升。

M3-CVC: Controllable Video Compression with Multimodal Generative Models

  • 用语义运动融合策略选关键帧,保留核心信息
  • 文本引导的扩散模型压缩关键帧,重建质量高
  • 适合需要精细控制和高质量还原的视频应用

传统与神经视频编码器在超低码率场景下普遍存在控制力弱、泛化性差的问题。为此,我们提出M3-CVC框架,结合多模态生成模型实现可控视频压缩。该框架采用语义-运动复合策略选取关键帧以保留关键信息;对每个关键帧及其对应视频片段,利用基于对话的大规模多模态模型(LMM)提取分层时空细节,增强帧间与帧内表征,提升视频保真度并改善编码可解释性。M3-CVC进一步采用基于条件扩散模型的文本引导关键帧压缩方法,在解码阶段,由LMM生成的文本描述指导扩散过程,精准恢复原始视频内容。实验表明,M3-CVC在超低码率场景下显著优于当前最先进的VVC标准,尤其在保持语义与感知保真度方面表现突出。

原文摘要 · Abstract (English)

Traditional and neural video codecs commonly encounter limitations in controllability and generality under ultra-low-bitrate coding scenarios. To overcome these challenges, we propose M3-CVC, a controllable video compression framework incorporating multimodal generative models. The framework utilizes a semantic-motion composite strategy for keyframe selection to retain critical information. For each keyframe and its corresponding video clip, a dialogue-based large multimodal model (LMM) approach extracts hierarchical spatiotemporal details, enabling both inter-frame and intra-frame representations for improved video fidelity while enhancing encoding interpretability. M3-CVC further employs a conditional diffusion-based, text-guided keyframe compression method, achieving high fidelity in frame reconstruction. During decoding, textual descriptions derived from LMMs guide the diffusion process to restore the original video's content accurately. Experimental results demonstrate that M3-CVC significantly outperforms the state-of-the-art VVC standard in ultra-low bitrate scenarios, particularly in preserving semantic and perceptual fidelity.

视频压缩多模态扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。