arXiv:2411.18335cs.CVcs.AI2024-11CVPR被引 7

构建首个真实世界全景立体深度数据集,推动全景深度估计发展

Helvipad: A Real-World Dataset for Omnidirectional Stereo Depth Estimation

  • 用双360°摄像头+激光雷达采集4万帧全景视频,生成精确深度标签
  • 在全景图像上基准测试主流立体模型,发现深度估计仍有明显不足
  • 提出适配方案提升模型性能,适合做全景视觉与自动驾驶研究者

尽管立体深度估计取得进展,全景成像仍因缺乏合适数据而研究不足。我们推出Helvipad,一个真实世界的全景立体深度估计数据集,包含来自多样环境(如拥挤室内外场景、不同光照条件)的4万帧视频序列。数据由上下布置的两台360°相机与激光雷达采集,通过将3D点云投影到等距柱状图生成准确的深度和视差标签。此外,我们通过深度补全构建了标签密度更高的增强训练集。我们在标准图像与全景图像上对领先立体深度估计模型进行基准测试,结果表明:尽管近期方法表现尚可,但全景成像中的深度估计仍面临挑战。为此,我们引入必要的模型适配,显著提升性能。

原文摘要 · Abstract (English)

Despite progress in stereo depth estimation, omnidirectional imaging remains underexplored, mainly due to the lack of appropriate data. We introduce Helvipad, a real-world dataset for omnidirectional stereo depth estimation, featuring 40K video frames from video sequences across diverse environments, including crowded indoor and outdoor scenes with various lighting conditions. Collected using two 360° cameras in a top-bottom setup and a LiDAR sensor, the dataset includes accurate depth and disparity labels by projecting 3D point clouds onto equirectangular images. Additionally, we provide an augmented training set with an increased label density by using depth completion. We benchmark leading stereo depth estimation models for both standard and omnidirectional images. The results show that while recent stereo methods perform decently, a challenge persists in accurately estimating depth in omnidirectional imaging. To address this, we introduce necessary adaptations to stereo models, leading to improved performance.

全景视觉深度估计数据集三维感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。