同时估计全景图的深度与法向量,提升复杂场景的3D理解能力。
Multi-task Geometric Estimation of Depth and Surface Normal from Monocular 360° Images
- 多任务学习架构融合深度与法向量信息,增强3D结构理解。
- 融合模块共享特征表示,显著提升复杂纹理下的预测精度。
- 在复杂多变场景中表现优越,适合全景视觉应用开发。
全景360°图像中的几何估计对场景理解至关重要。现有方法通常仅预测单一特征(如深度或表面法向量),在处理复杂纹理或曲面时鲁棒性不足。本文提出一种新型多任务学习(MTL)网络,可同时从360°图像中估计深度与表面法向量。首先,设计了一种新颖的MTL架构,通过整合深度与法向量的几何信息,提升两任务的预测性能,实现更深入的3D场景结构理解。其次,引入融合模块连接两个任务,使网络能够学习共享表征,从而提高准确性和鲁棒性。实验表明,该模型在深度与表面法向量估计上均显著优于现有最优方法,在复杂多变场景中表现突出。其有效性与泛化能力使其成为360°图像几何估计的新基准。代码与模型已开源:https://github.com/huangkun101230/360MTLGeometricEstimation。
原文摘要 · Abstract (English)
Geometric estimation is required for scene understanding and analysis in panoramic 360° images. Current methods usually predict a single feature, such as depth or surface normal. These methods can lack robustness, especially when dealing with intricate textures or complex object surfaces. We introduce a novel multi-task learning (MTL) network that simultaneously estimates depth and surface normals from 360° images. Our first innovation is our MTL architecture, which enhances predictions for both tasks by integrating geometric information from depth and surface normal estimation, enabling a deeper understanding of 3D scene structure. Another innovation is our fusion module, which bridges the two tasks, allowing the network to learn shared representations that improve accuracy and robustness. Experimental results demonstrate that our MTL architecture significantly outperforms state-of-the-art methods in both depth and surface normal estimation, showing superior performance in complex and diverse scenes. Our model's effectiveness and generalizability, particularly in handling intricate surface textures, establish it as a new benchmark in 360° image geometric estimation. The code and model are available at \url{https://github.com/huangkun101230/360MTLGeometricEstimation}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。