提出双向视频压缩新框架,显著降低码率并超越现有标准
BiECVC: Gated Diversification of Bidirectional Contexts for Learned Video Compression
- 通过双向上下文建模与自适应门控机制提升预测准确性
- 在随机访问配置下比VTM 13.2降低13.4%~15.7%码率
- 首次在所有标准数据集上超越VTM 13.2的学界视频编码方法
近期基于前向预测的可学习视频压缩(LVC)方法已取得优异成果,甚至在低延迟B(LDB)配置下超越VVC参考软件VTM。相比之下,可学习双向视频压缩(BVC)研究不足,性能仍落后于单向方法。其主要原因是难以提取多样且准确的上下文:现有BVC主要依赖时间运动信息,忽视帧间非局部相关性,且缺乏对快速运动或遮挡引发有害上下文的动态抑制能力。为此,本文提出BiECVC框架,融合多样化局部与非局部上下文建模及自适应上下文门控。局部上下文增强方面,复用低层高质量特征并利用解码运动矢量对齐,无额外运动开销;为高效建模非局部依赖,采用线性注意力机制平衡性能与复杂度;为缓解不准确上下文预测的影响,引入受自回归语言模型中数据依赖衰减启发的双向上下文门控,根据条件编码结果动态过滤上下文信息。大量实验表明,BiECVC在随机访问(RA)配置下,分别在内帧周期为32和64时,比特率较VTM 13.2降低13.4%和15.7%。据我们所知,BiECVC是首个在所有标准测试数据集上均超越VTM 13.2 RA的学界视频编解码器。
原文摘要 · Abstract (English)
Recent forward prediction-based learned video compression (LVC) methods have achieved impressive results, even surpassing VVC reference software VTM under the Low Delay B (LDB) configuration. In contrast, learned bidirectional video compression (BVC) remains underexplored and still lags behind its forward-only counterparts. This performance gap is mainly due to the limited ability to extract diverse and accurate contexts: most existing BVCs primarily exploit temporal motion while neglecting non-local correlations across frames. Moreover, they lack the adaptability to dynamically suppress harmful contexts arising from fast motion or occlusion. To tackle these challenges, we propose BiECVC, a BVC framework that incorporates diversified local and non-local context modeling along with adaptive context gating. For local context enhancement, BiECVC reuses high-quality features from lower layers and aligns them using decoded motion vectors without introducing extra motion overhead. To model non-local dependencies efficiently, we adopt a linear attention mechanism that balances performance and complexity. To further mitigate the impact of inaccurate context prediction, we introduce Bidirectional Context Gating, inspired by data-dependent decay in recent autoregressive language models, to dynamically filter contextual information based on conditional coding results. Extensive experiments demonstrate that BiECVC achieves state-of-the-art performance, reducing the bit-rate by 13.4% and 15.7% compared to VTM 13.2 under the Random Access (RA) configuration with intra periods of 32 and 64, respectively. To our knowledge, BiECVC is the first learned video codec to surpass VTM 13.2 RA across all standard test datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。