arXiv:2509.03922cs.CV2025-09

提升多视角视频压缩效率,支持灵活视角切换且兼容单视图解码。

DCVC-MV: Deep Contextual Multiview Video Compression with Efficient Inter-View Prediction

  • 通过视角间运动特征传播与条件熵模型,提升多视角关联建模能力。
  • 在MVD、LLVIP等数据集上实现比特率降低15%以上,压缩性能显著提升。
  • 适合虚拟现实、自由视角直播等需高效传输多视角视频的场景。

多视角视频是自由视角广播和虚拟现实等3D应用的关键格式,但其庞大的数据量给存储与传输带来巨大挑战。随着深度上下文视频压缩技术逐步成熟并迈向标准化,将此类学习型编码器扩展至多视角场景成为实际部署的必要方向,但该领域仍缺乏系统探索。本文提出DCVC-MV,一种新型深度上下文多视角视频压缩框架,满足三项核心要求:保持向后兼容性,使主视角码流可独立由单视角解码器解码;支持随机访问能力,实现不同视角间的灵活切换;有效利用视角间相关性以提升压缩效率。该框架包含四个专用组件:(1) 视角间运动特征传播方法,将已解码独立视角的运动特征作为条件传递至依赖视角以促进运动编码;(2) 视角间运动条件熵模型,用于跨视角学习运动潜在表示的概率先验,实现更精确的概率估计;(3) 隐式视角间上下文预测方法,从低分辨率独立视角内容特征中预测视角间上下文,无需显式视差估计;(4) 视角间上下文条件熵模型,学习跨视角的上下文条件先验,进一步增强内容压缩效果。

原文摘要 · Abstract (English)

Multiview video is a key format for 3D applications such as free-viewpoint broadcasting and virtual reality, yet its large data volume poses significant challenges for efficient storage and transmission. As deep contextual video compression matures and moves toward standardization, extending such learned codecs to multiview scenarios has become essential for practical deployment---yet this direction remains largely unexplored. In this paper, we propose DCVC-MV, a novel deep contextual multiview video compression framework that satisfies three fundamental requirements. First, it maintains backward compatibility, ensuring that the primary view's bitstream can be decoded independently by a single-view decoder without being affected by other views. Second, it supports random-access capability, enabling flexible switching between different viewing perspectives. Third, it effectively exploits inter-view correlations to achieve high compression efficiency. This is realized through four dedicated components: (1) an inter-view motion feature propagation method, which propagates decoded independent-view motion features as conditions to promote dependent-view motion encoding; (2) an inter-view motion conditional entropy model designed to learn motion conditional priors across views for more accurate probability estimation of motion latent representations; (3) an implicit inter-view context prediction method, which predicts inter-view contexts from low-resolution independent-view content features without explicit disparity estimation; and (4) an inter-view contextual conditional entropy model that learns contextual conditional priors across views to further enhance content compression.

多视角视频深度压缩上下文建模虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。