arXiv:2409.04013cs.CVcs.IT2024-09被引 1

用3D高斯点云提升多视角图像压缩精度

3D-LMVIC: Learning-based Multi-View Image Coding with 3D Gaussian Geometric Priors

  • 基于3D高斯点云构建几何先验,精准估计视差
  • 相比传统方法,压缩性能提升显著,视差估计更准
  • 适合虚拟现实、自动驾驶等宽基线多摄像头场景

现有多视角图像压缩方法通常依赖视图间的2D投影相似性来估计视差。虽然在小视差(如立体图像)下有效,但在虚拟现实与自动驾驶中常见的宽基线多摄像机系统所面临的复杂视差面前表现不佳。为此,我们提出3D-LMVIC,一种基于学习的多视角图像压缩框架,利用3D高斯点云渲染生成几何先验,实现更准确的视差估计。同时,引入深度图压缩模型以减少视图间几何冗余,并设计基于视图间距离度量的多视图序列排序策略,增强相邻视图相关性。实验表明,3D-LMVIC在性能上优于传统及基于学习的方法,且显著提升两视图方法的视差估计精度。

原文摘要 · Abstract (English)

Existing multi-view image compression methods often rely on 2D projection-based similarities between views to estimate disparities. While effective for small disparities, such as those in stereo images, these methods struggle with the more complex disparities encountered in wide-baseline multi-camera systems, commonly found in virtual reality and autonomous driving applications. To address this limitation, we propose 3D-LMVIC, a novel learning-based multi-view image compression framework that leverages 3D Gaussian Splatting to derive geometric priors for accurate disparity estimation. Furthermore, we introduce a depth map compression model to minimize geometric redundancy across views, along with a multi-view sequence ordering strategy based on a defined distance measure between views to enhance correlations between adjacent views. Experimental results demonstrate that 3D-LMVIC achieves superior performance compared to both traditional and learning-based methods. Additionally, it significantly improves disparity estimation accuracy over existing two-view approaches.

多视角压缩3D高斯视差估计虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。