arXiv:2512.20871cs.CVcs.MM2025-12被引 1

只解码用户视线范围,大幅提升360度视频压缩效率

NeRV360: Neural Representation for 360-Degree Videos with a Viewport Decoder

  • 仅解码用户视角区域,避免全图重建
  • 内存减少7倍,解码速度提升2.5倍
  • 适合实时360度视频应用的轻量级方案

针对高分辨率360度视频,隐式神经视频表示(NeRV)面临内存占用高、解码慢的问题。本文提出NeRV360,一种端到端框架,仅解码用户选定的视口而非整个全景帧。与传统流程不同,NeRV360将视口提取融入解码过程,并引入时空自适应仿射变换模块,实现基于视角和时间的条件解码。在6K分辨率视频上的实验表明,相较于代表性方法HNeRV,NeRV360在内存消耗上降低7倍,解码速度提升2.5倍,且在客观指标上表现更优。

原文摘要 · Abstract (English)

Implicit neural representations for videos (NeRV) have shown strong potential for video compression. However, applying NeRV to high-resolution 360-degree videos causes high memory usage and slow decoding, making real-time applications impractical. We propose NeRV360, an end-to-end framework that decodes only the user-selected viewport instead of reconstructing the entire panoramic frame. Unlike conventional pipelines, NeRV360 integrates viewport extraction into decoding and introduces a spatial-temporal affine transform module for conditional decoding based on viewpoint and time. Experiments on 6K-resolution videos show that NeRV360 achieves a 7-fold reduction in memory consumption and a 2.5-fold increase in decoding speed compared to HNeRV, a representative prior work, while delivering better image quality in terms of objective metrics.

360视频神经表示视口解码高效压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。