只解码用户视线范围,大幅提升360度视频压缩效率
NeRV360: Neural Representation for 360-Degree Videos with a Viewport Decoder
- 仅解码用户视角区域,避免全图重建
- 内存减少7倍,解码速度提升2.5倍
- 适合实时360度视频应用的轻量级方案
针对高分辨率360度视频,隐式神经视频表示(NeRV)面临内存占用高、解码慢的问题。本文提出NeRV360,一种端到端框架,仅解码用户选定的视口而非整个全景帧。与传统流程不同,NeRV360将视口提取融入解码过程,并引入时空自适应仿射变换模块,实现基于视角和时间的条件解码。在6K分辨率视频上的实验表明,相较于代表性方法HNeRV,NeRV360在内存消耗上降低7倍,解码速度提升2.5倍,且在客观指标上表现更优。
原文摘要 · Abstract (English)
Implicit neural representations for videos (NeRV) have shown strong potential for video compression. However, applying NeRV to high-resolution 360-degree videos causes high memory usage and slow decoding, making real-time applications impractical. We propose NeRV360, an end-to-end framework that decodes only the user-selected viewport instead of reconstructing the entire panoramic frame. Unlike conventional pipelines, NeRV360 integrates viewport extraction into decoding and introduces a spatial-temporal affine transform module for conditional decoding based on viewpoint and time. Experiments on 6K-resolution videos show that NeRV360 achieves a 7-fold reduction in memory consumption and a 2.5-fold increase in decoding speed compared to HNeRV, a representative prior work, while delivering better image quality in terms of objective metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。