提升视频压缩质量,通过长时空上下文增强实现更高画质与更低码率。
L-STEC: Learned Video Compression with Long-term Spatio-Temporal Enhanced Context
- 引入LSTM扩展参考帧链,捕捉长期时序依赖关系。
- 融合像素域的形变空间上下文,提升细节保留能力。
- 在多个指标上超越现有方法,适合高效视频编码研究者。
近年来,神经视频压缩技术兴起,条件化框架已超越传统编解码器。然而,多数方法仅依赖前一帧特征预测时间上下文,导致两大问题:一是短参考窗口遗漏长期依赖和细微纹理;二是仅传播特征级信息会累积误差,造成预测不准与细节丢失。为此,本文提出长时空增强上下文(L-STEC)方法:首先利用LSTM扩展参考链以捕捉长期依赖;再从像素域引入形变空间上下文,通过多感受野网络融合时空信息,更好地保留参考细节。实验表明,L-STEC显著提升压缩性能,在PSNR上相比DCVC-TCM节省37.01%码率,在MS-SSIM上节省31.65%,优于VTM-17.0与DCVC-FM,达到当前最佳水平。
原文摘要 · Abstract (English)
Neural Video Compression has emerged in recent years, with condition-based frameworks outperforming traditional codecs. However, most existing methods rely solely on the previous frame's features to predict temporal context, leading to two critical issues. First, the short reference window misses long-term dependencies and fine texture details. Second, propagating only feature-level information accumulates errors over frames, causing prediction inaccuracies and loss of subtle textures. To address these, we propose the Long-term Spatio-Temporal Enhanced Context (L-STEC) method. We first extend the reference chain with LSTM to capture long-term dependencies. We then incorporate warped spatial context from the pixel domain, fusing spatio-temporal information through a multi-receptive field network to better preserve reference details. Experimental results show that L-STEC significantly improves compression by enriching contextual information, achieving 37.01% bitrate savings in PSNR and 31.65% in MS-SSIM compared to DCVC-TCM, outperforming both VTM-17.0 and DCVC-FM and establishing new state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。