arXiv:2410.05735cs.CV2024-10被引 2

用立方体场模型从单张全景图连续估计任意视角深度,提升VR沉浸感。

CUBE360: Learning Cubic Field Representation for Monocular 360 Depth Estimation for Virtual Reality

  • 将全景图转为六面立方图,用多平面图像构建立方体场表示。
  • 自监督训练下在合成与真实数据上均优于现有方法,支持任意视角渲染。
  • 适合需要高沉浸感的VR漫游与视觉特效应用,无需深度标注。

全景图提供全面场景信息,适用于虚拟现实(VR)应用。获取对应深度图对实现沉浸式交互体验至关重要。然而,由于等距柱状投影(ERP)带来的严重畸变以及全景RGB-D数据集稀缺,全景深度估计面临巨大挑战。受神经渲染成功启发,我们提出一种新方法CUBE360,从单张全景图学习由多个多平面图像(MPIs)组成的立方体场,实现任意视角下的连续深度估计。该方法采用立方图投影将ERP图像转换为六个面并分别提取各面的MPI,降低高分辨率数据处理所需内存,同时避免了ERP中像素分布不均带来的计算复杂性。进一步引入基于注意力的融合模块,学习立方体各面MPI间的关联,构建包含颜色与密度信息的分层立方体场表示。此外,提出一种新型采样策略,可在立方体与平面尺度上从立方体场渲染新视角。整个流程通过自监督学习框架中的光度损失进行训练,可直接使用无深度标注的360视频。在合成与真实数据集上的实验表明,CUBE360性能显著优于现有自监督方法。同时,其在VR漫游与视觉特效等下游任务中也展现出高效性,凸显其增强沉浸体验的潜力。

原文摘要 · Abstract (English)

Panoramic images provide comprehensive scene information and are suitable for VR applications. Obtaining corresponding depth maps is essential for achieving immersive and interactive experiences. However, panoramic depth estimation presents significant challenges due to the severe distortion caused by equirectangular projection (ERP) and the limited availability of panoramic RGB-D datasets. Inspired by the recent success of neural rendering, we propose a novel method, named $\mathbf{CUBE360}$, that learns a cubic field composed of multiple MPIs from a single panoramic image for $\mathbf{continuous}$ depth estimation at any view direction. Our CUBE360 employs cubemap projection to transform an ERP image into six faces and extract the MPIs for each, thereby reducing the memory consumption required for MPI processing of high-resolution data. Additionally, this approach avoids the computational complexity of handling the uneven pixel distribution inherent to equirectangular projectio. An attention-based blending module is then employed to learn correlations among the MPIs of cubic faces, constructing a cubic field representation with color and density information at various depth levels. Furthermore, a novel sampling strategy is introduced for rendering novel views from the cubic field at both cubic and planar scales. The entire pipeline is trained using photometric loss calculated from rendered views within a self-supervised learning approach, enabling training on 360 videos without depth annotations. Experiments on both synthetic and real-world datasets demonstrate the superior performance of CUBE360 compared to prior SSL methods. We also highlight its effectiveness in downstream applications, such as VR roaming and visual effects, underscoring CUBE360's potential to enhance immersive experiences.

全景深度虚拟现实立方体场自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。