通过融合空间与时间信息,提升视频编码对大运动和新物体的处理能力。
Augmented Deep Contexts for Spatially Embedded Video Coding
- 结合空间与时间参考生成增强的运动向量和混合上下文
- 引入空间引导的潜在先验,减少潜在表示错位问题
- 联合优化空间-时间码率分配,降低11.9%比特率
现有神经视频编码器(NVC)仅依赖时间参考生成时序上下文和潜在先验,导致在处理大运动或新出现物体时受限于上下文不足和潜在先验错位。为此,我们提出空间嵌入式视频编码器(SEVC),将低分辨率视频压缩以提供空间参考。首先,SEVC同时利用空间与时间参考生成增强的运动向量和混合空间-时序上下文;其次,为解决潜在先验错位问题并丰富先验信息,引入由多个时序潜在表示增强的空间引导潜在先验;最后,设计联合空间-时序优化策略,学习质量自适应的码率分配,进一步提升率失真性能。实验表明,所提SEVC有效缓解了大运动与新物体处理难题,在保持更高画质的同时,比特率比当前最优方法降低11.9%,且额外输出低分辨率码流。代码与模型已开源于https://github.com/EsakaK/SEVC。
原文摘要 · Abstract (English)
Most Neural Video Codecs (NVCs) only employ temporal references to generate temporal-only contexts and latent prior. These temporal-only NVCs fail to handle large motions or emerging objects due to limited contexts and misaligned latent prior. To relieve the limitations, we propose a Spatially Embedded Video Codec (SEVC), in which the low-resolution video is compressed for spatial references. Firstly, our SEVC leverages both spatial and temporal references to generate augmented motion vectors and hybrid spatial-temporal contexts. Secondly, to address the misalignment issue in latent prior and enrich the prior information, we introduce a spatial-guided latent prior augmented by multiple temporal latent representations. At last, we design a joint spatial-temporal optimization to learn quality-adaptive bit allocation for spatial references, further boosting rate-distortion performance. Experimental results show that our SEVC effectively alleviates the limitations in handling large motions or emerging objects, and also reduces 11.9% more bitrate than the previous state-of-the-art NVC while providing an additional low-resolution bitstream. Our code and model are available at https://github.com/EsakaK/SEVC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。