提升长距离视频运动估计与预测,实现更高效的自学习双向压缩
L-LBVC: Long-Term Motion Estimation and Prediction for Learned Bi-Directional Video Compression
- 分层处理短距与长距运动,递归累积局部光流估算远距离运动
- 测试时自适应降采样参考帧,匹配训练时的运动范围,降低编码开销
- 在随机访问下超越VVC,在复杂运动场景中表现突出
近年来,自学习视频压缩(LVC)在低延迟配置下表现出色。然而,自学习双向视频压缩(LBVC)的性能仍落后于传统双向编码,主要源于对远距离帧的长期运动估计与预测不准确,尤其在大运动场景中。本文提出新型LBVC框架L-LBVC:首先设计自适应运动估计模块,直接估计相邻帧及小运动非相邻帧的光流;对大运动非相邻帧,通过递归累积相邻帧间的局部光流来估算长期光流。其次提出自适应运动预测模块,显著降低运动编码的比特开销。为提升长期运动预测精度,测试时自适应降采样参考帧以匹配训练时观察到的运动范围。实验表明,所提L-LBVC显著优于先前最先进的LVC方法,甚至在部分测试数据集上超越VVC(VTM)在随机访问配置下的表现。
原文摘要 · Abstract (English)
Recently, learned video compression (LVC) has shown superior performance under low-delay configuration. However, the performance of learned bi-directional video compression (LBVC) still lags behind traditional bi-directional coding. The performance gap mainly arises from inaccurate long-term motion estimation and prediction of distant frames, especially in large motion scenes. To solve these two critical problems, this paper proposes a novel LBVC framework, namely L-LBVC. Firstly, we propose an adaptive motion estimation module that can handle both short-term and long-term motions. Specifically, we directly estimate the optical flows for adjacent frames and non-adjacent frames with small motions. For non-adjacent frames with large motions, we recursively accumulate local flows between adjacent frames to estimate long-term flows. Secondly, we propose an adaptive motion prediction module that can largely reduce the bit cost for motion coding. To improve the accuracy of long-term motion prediction, we adaptively downsample reference frames during testing to match the motion ranges observed during training. Experiments show that our L-LBVC significantly outperforms previous state-of-the-art LVC methods and even surpasses VVC (VTM) on some test datasets under random access configuration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。