arXiv:2510.15775eess.IVcs.CV2025-10

提出场景感知神经表示框架,实现光场图像高效压缩。

SANR: Scene-Aware Neural Representation for Light Field Image Compression with Rate-Distortion Optimization

  • 引入分层场景建模块,用多尺度隐变量捕捉场景结构。
  • 首次在神经表示中引入熵约束量化训练,实现端到端率失真优化。
  • 相比HEVC提升65.62%率失真性能,适合3D重建与传输场景。

光场图像包含多视角场景信息,在三维重建中至关重要。但其高维特性导致数据量巨大,给实际存储与传输带来挑战。尽管基于神经表示的方法在光场图像压缩中展现出潜力,但多数依赖隐式神经表示的坐标到像素直接映射,忽视了显式场景结构建模,且缺乏端到端率失真优化,限制了压缩效率。为此,本文提出SANR:一种具有端到端率失真优化能力的场景感知神经表示框架。为增强场景感知,SANR设计分层场景建模模块,利用多尺度隐码捕捉内在场景结构,缩小输入坐标与目标光场图像之间的信息差距。从压缩角度,SANR首次将熵约束量化感知训练(QAT)引入基于神经表示的光场图像压缩,实现端到端率失真优化。大量实验表明,SANR在率失真性能上显著优于现有技术,相较HEVC实现65.62%的BD-rate降低。

原文摘要 · Abstract (English)

Light field images capture multi-view scene information and play a crucial role in 3D scene reconstruction. However, their high-dimensional nature results in enormous data volumes, posing a significant challenge for efficient compression in practical storage and transmission scenarios. Although neural representation-based methods have shown promise in light field image compression, most approaches rely on direct coordinate-to-pixel mapping through implicit neural representation (INR), often neglecting the explicit modeling of scene structure. Moreover, they typically lack end-to-end rate-distortion optimization, limiting their compression efficiency. To address these limitations, we propose SANR, a Scene-Aware Neural Representation framework for light field image compression with end-to-end rate-distortion optimization. For scene awareness, SANR introduces a hierarchical scene modeling block that leverages multi-scale latent codes to capture intrinsic scene structures, thereby reducing the information gap between INR input coordinates and the target light field image. From a compression perspective, SANR is the first to incorporate entropy-constrained quantization-aware training (QAT) into neural representation-based light field image compression, enabling end-to-end rate-distortion optimization. Extensive experiment results demonstrate that SANR significantly outperforms state-of-the-art techniques regarding rate-distortion performance with a 65.62\% BD-rate saving against HEVC.

光场图像神经表示率失真优化压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。