arXiv:2506.07709eess.IVcs.CV2025-06被引 3

提出细粒度运动压缩与选择性时序融合,提升神经B帧视频编码效率。

Fine-Grained Motion Compression and Selective Temporal Fusion for Neural B-Frame Video Coding

  • 设计双分支交互式运动自编码器,实现双向运动向量的精细压缩。
  • 在随机访问配置下,相比顶尖神经B帧编码器降低约10%的比特率。
  • 适合关注视频编码效率与低延迟场景的研究者与工程师。

随着神经P帧视频编码的显著进展,神经B帧编码已成为关键研究方向。然而,现有大多数神经B帧编解码器直接沿用P帧工具,未充分应对B帧压缩的独特挑战,导致性能不佳。为此,本文提出针对神经B帧编码的运动压缩与时序融合新方法。首先,设计细粒度运动压缩机制:采用交互式双分支运动自编码器,结合分支自适应量化步长,实现双向运动向量的精细化压缩,并满足其不对称码率分配与重建质量需求;同时引入交互式运动熵模型,通过分区潜在表示作为方向先验,挖掘双向运动潜在表征间的相关性。其次,提出选择性时序融合方法:预测双向融合权重,实现对多尺度时序上下文的差异化利用;并引入基于超先验的隐式对齐机制,将超先验作为上下文潜在表示的代理,隐式缓解融合后双向时序先验的错位问题。大量实验表明,所提编解码器在平均BD-rate上相比当前最优神经B帧编码器DCVC-B降低约10%,且在随机访问配置下达到甚至超越H.266/VVC参考软件的压缩性能。

原文摘要 · Abstract (English)

With the remarkable progress in neural P-frame video coding, neural B-frame coding has recently emerged as a critical research direction. However, most existing neural B-frame codecs directly adopt P-frame coding tools without adequately addressing the unique challenges of B-frame compression, leading to suboptimal performance. To bridge this gap, we propose novel enhancements for motion compression and temporal fusion for neural B-frame coding. First, we design a fine-grained motion compression method. This method incorporates an interactive dual-branch motion auto-encoder with per-branch adaptive quantization steps, which enables fine-grained compression of bi-directional motion vectors while accommodating their asymmetric bitrate allocation and reconstruction quality requirements. Furthermore, this method involves an interactive motion entropy model that exploits correlations between bi-directional motion latent representations by interactively leveraging partitioned latent segments as directional priors. Second, we propose a selective temporal fusion method that predicts bi-directional fusion weights to achieve discriminative utilization of bi-directional multi-scale temporal contexts with varying qualities. Additionally, this method introduces a hyperprior-based implicit alignment mechanism for contextual entropy modeling. By treating the hyperprior as a surrogate for the contextual latent representation, this mechanism implicitly mitigates the misalignment in the fused bi-directional temporal priors. Extensive experiments demonstrate that our proposed codec achieves an average BD-rate reduction of approximately 10% compared to the state-of-the-art neural B-frame codec, DCVC-B, and delivers comparable or even superior compression performance to the H.266/VVC reference software under random-access configurations.

视频编码神经B帧运动压缩时序融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。