arXiv:2509.16706eess.IV2025-09被引 2

用多网格隐式表示压缩多视角视频,大幅降低码率

A Multi-Grid Implicit Neural Representation for Multi-View Videos

  • 设计时间、视角与融合三重网格捕捉共性与细节
  • 相比TMIV标准,码率降低72.3%且保持同等画质
  • 适合高分辨率多视角视频的高效编码与传输

多视角视频在多个领域广泛应用,但其高分辨率和多相机拍摄带来存储与传输挑战。本文提出MV-MGINR,一种用于多视角视频的多网格隐式神经表示。该方法结合时间索引网格、视角索引网格以及时空融合网格,分别捕捉各视角的共性内容、时间轴上的共性特征以及特定视角与时间下的局部细节。随后通过合成网络上采样多网格潜在表示,生成重建帧。同时引入运动感知损失,提升运动区域的重建质量。该框架有效整合多视角视频的共性与局部特征,最终实现高质量重建。相较于MPEG沉浸式视频测试模型TMIV,MV-MGINR在保持相同峰值信噪比(PSNR)的前提下,实现72.3%的码率节省。

原文摘要 · Abstract (English)

Multi-view videos are becoming widely used in different fields, but their high resolution and multi-camera shooting raise significant challenges for storage and transmission. In this paper, we propose MV-MGINR, a multi-grid implicit neural representation for multi-view videos. It combines a time-indexed grid, a view-indexed grid and an integrated time and view grid. The first two grids capture common representative contents across each view and time axis respectively, and the latter one captures local details under specific view and time. Then, a synthesis net is used to upsample the multi-grid latents and generate reconstructed frames. Additionally, a motion-aware loss is introduced to enhance the reconstruction quality of moving regions. The proposed framework effectively integrates the common and local features of multi-view videos, ultimately achieving high-quality reconstruction. Compared with MPEG immersive video test model TMIV, MV-MGINR achieves bitrate savings of 72.3% while maintaining the same PSNR.

多视角视频隐式表示码率压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。