arXiv:2410.09706eess.IV2024-10CVPR被引 32

利用多帧非局部相关性提升视频压缩效率,减少误差累积。

ECVC: Exploiting Non-Local Correlations in Multiple Frames for Contextual Video Compression

  • 通过多帧非局部相关建模增强时间上下文预测能力。
  • 在IP=32和IP=-1下分别比SOTA方法少10.5%和11.5%码率。
  • 适合追求高码率效率的视频压缩研究与应用开发者。

在学习型视频压缩(LVC)中,提升帧间预测能力,如增强时间上下文挖掘和缓解误差累积,对改善率失真性能至关重要。现有LVC主要关注时序运动建模,忽视了帧间的非局部相关性。此外,当前上下文视频压缩模型仅使用单个参考帧,难以处理复杂运动。为此,我们提出利用多帧间的非局部相关性来增强时间先验,显著提升率失真性能。为缓解误差累积,引入部分级联微调策略,在有限计算资源下支持长序列微调,降低训练与测试序列长度不匹配问题,显著减少误差累积。基于上述技术,我们提出了视频压缩方案ECVC。实验表明,相较于此前SOTA方法DCVC-FM,ECVC在VTM-13.2低延迟B模式下,于内插周期(IP)为32和-1时,分别减少10.5%和11.5%的比特率。

原文摘要 · Abstract (English)

In Learned Video Compression (LVC), improving inter prediction, such as enhancing temporal context mining and mitigating accumulated errors, is crucial for boosting rate-distortion performance. Existing LVCs mainly focus on mining the temporal movements while neglecting non-local correlations among frames. Additionally, current contextual video compression models use a single reference frame, which is insufficient for handling complex movements. To address these issues, we propose leveraging non-local correlations across multiple frames to enhance temporal priors, significantly boosting rate-distortion performance. To mitigate error accumulation, we introduce a partial cascaded fine-tuning strategy that supports fine-tuning on full-length sequences with constrained computational resources. This method reduces the train-test mismatch in sequence lengths and significantly decreases accumulated errors. Based on the proposed techniques, we present a video compression scheme ECVC. Experiments demonstrate that our ECVC achieves state-of-the-art performance, reducing 10.5% and 11.5% more bit-rates than previous SOTA method DCVC-FM over VTM-13.2 low delay B (LDB) under the intra period (IP) of 32 and -1, respectively.

视频压缩非局部相关误差累积LVC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。