用状态空间模型提升神经视频编码的双向预测能力,实现低延迟高效率压缩。
DCVC-MB: Neural B-Frame Video Compression using State Space Models

- 引入状态空间模型构建时空融合网络,实现双向时间预测。
- 在多个基准上实现最高8.98%的BD-rate降低,优于现有神经编码器。
- 适合追求高效视频压缩与低延迟应用的研究者与工程师。
本文提出一种面向B帧编码的神经视频编解码框架DCVC-Mamba(DCVC-MB)。该方法采用低延迟B帧编码的IBP帧策略,基于状态空间模型构建时空融合网络以实现双向时间预测,并设计了感知熵的跳过机制,选择性跳过部分潜在表示以减少熵编码时间。此外,还实现了两种推理时优化策略以进一步提升压缩性能。实验表明,相较于现有神经视频编解码器,DCVC-MB平均实现8.98%的BD-rate降低;相比VTM-19.0-LDP和VTM-19.0-RA(Inter-GoP=16)基准,分别提升30.45%和1.81%,推动了神经视频压缩技术的发展。
原文摘要 · Abstract (English)
In this paper we propose DCVC-Mamba (DCVC-MB), a neural video codec framework for B-frame coding. Our approach incorporates an IBP frame strategy for low-delay B-frame coding, a spatio-temporal fusion model based on state-space models for bidirectional temporal prediction, and an entropy-aware skipping mechanism that selectively omits coding certain latents to reduce entropy coding times. In addition to our model contributions we also implement two inference-time strategies that enhance compression performance. Experimental evaluation shows that DCVC-MB compares favorably to existing NVCs and traditional codecs. The method demonstrates BD-rate reductions of up to $8.98\%$ on average compared to prior neural video codecs, and improvements of up to $30.45\%$ and $1.81\%$ over the VTM-19.0-LDP and VTM-19.0-RA(Inter-GoP=16) benchmarks, respectively, contributing to advances in neural video compression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。