arXiv:2608.28689cs.CV2026-08

不同投影方式对360度视频神经压缩效率影响显著,需根据编码器类型选择最优投影。

Projection-Aware End-to-End Learned Video Compression for 360-Degree Video

论文配图:Projection-Aware End-to-End Learned Video Compression for 360-Degree Video
图 1 · 摘自论文原文
  • 构建可微分流水线,联合优化投影转换、神经压缩与逆投影过程。
  • 等距柱状图和填充等距柱状图在神经模型下压缩效率最高,优于立方体贴图等格式。
  • 投影选择依赖编码器类型:光流模型偏好单面连续投影,块编码更适应多面布局。

360度视频支持虚拟现实、自动驾驶和教育等沉浸式应用。由于球面内容无法被传统视频编解码器直接处理,必须先映射到二维投影。投影选择影响空间连续性、采样均匀性、运动估计和压缩效率。本文研究投影格式对360度视频端到端神经压缩的影响。采用JVET 360Lib支持的七种投影格式,基于尺度空间光流模型、JVET测试序列和通用测试条件进行评估。每个序列从源等距柱状图转换为编码投影,于多个码率点压缩、重建并反向转换。性能通过PSNR、球面PSNR、加权球面PSNR及Bjøntegaard delta rate评估。同时对比了结合投影转换、神经压缩与逆投影的可微分流水线与360Lib方法。结果表明,等距柱状图和填充等距柱状图在尺度空间光流模型下表现最佳,而基于立方体和菱形十二面体的投影效果较差。这与传统HM-16.16编码器相反,后者中立方体类投影(尤其是等角和调整立方体)优于等距柱状图。基于光流的神经模型受益于单面投影的空间连续性,而基于块的混合编码器则更适合多面布局。研究显示投影效率依赖于编解码器类型,为学习型360度视频压缩的投影选择提供指导。

原文摘要 · Abstract (English)

360-degree video supports immersive applications such as virtual reality, autonomous driving, and education. Because spherical content cannot be processed directly by conventional video codecs, it must first be mapped to a two-dimensional projection. Projection choice affects spatial continuity, sampling uniformity, motion estimation, and compression efficiency. This thesis investigates how projection format influences end-to-end neural compression of 360-degree video. Seven formats supported by JVET 360Lib are evaluated using the scale-space flow model, JVET test sequences, and common test conditions. Each sequence is converted from its source equirectangular projection to a coding projection, compressed at multiple rate points, reconstructed, and converted back. Performance is assessed using PSNR, spherical PSNR, weighted spherical PSNR, and Bjøntegaard delta rate. A differentiable pipeline combining projection conversion, neural compression, and inverse projection is also compared with 360Lib. Results show that equirectangular and padded equirectangular projections provide the highest compression efficiency with the scale-space flow model, while cubemap-based and rhombic dodecahedron projections are less effective. This differs from the conventional HM-16.16 codec, for which cubemap-based formats, particularly equi-angular and adjusted cubemap projections, outperform equirectangular formats. Neural models based on optical flow benefit from the spatial continuity of single-face projections, whereas block-based hybrid codecs better accommodate multi-face layouts. These findings show that projection efficiency is codec-dependent and provide guidance for selecting projections for learning-based 360-degree video compression.

360度视频神经压缩投影选择可微分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。