首个鱼眼深度估计高精度基准,支持毫米级真实值评估
WideDepth: Millimeter-Accurate Benchmark for Fisheye Depth Estimation

- 构建101个场景5000对高分辨率鱼眼图像,带毫米级真值
- 提出适配针孔模型到鱼眼的迁移方法,性能提升62%
- 适合机器人近场感知与深度估计研究者使用
鱼眼相机在机器人近场操作、导航和沉浸式感知中日益普及,但缺乏室内高精度深度真值基准。为此,我们提出WideDepth——首个面向鱼眼深度估计的室内数据集,包含101个场景、5000对高分辨率立体图像,标注有毫米级真实深度与视差。数据集还包含不同视场角与基线下的针孔与鱼眼配对样本,涵盖水平与垂直立体配置。我们提出一种将针孔训练的立体模型适配至鱼眼图像的方法,并基于高分辨率激光雷达扫描构建新颖的鱼眼立体图像生成流程。利用这些方法,我们在该基准上全面评估了前沿单目深度、立体匹配与深度补全模型。此外,提供18000个由激光雷达生成的稀疏深度训练样本,使针孔基础立体模型在鱼眼数据上性能最高提升62%。整体而言,本基准的高精度与多样性为鱼眼深度估计与机器人感知研究奠定了坚实基础。
原文摘要 · Abstract (English)
Fisheye cameras are increasingly adopted in robotics for near-field manipulation, navigation, and immersive perception, yet indoor depth benchmarks with accurate ground truth are still missing. To address this, we introduce WideDepth - the first indoor dataset for fisheye depth estimation, featuring 101 scenes containing 5K high-resolution stereo pairs labeled with millimeter-level ground truth depth and disparity. Our dataset also includes paired pinhole and fisheye samples across varying fields of view and baselines in both horizontal and vertical stereo setups. We further propose a method to adapt pinhole-trained stereo models to fisheye images and introduce a novel stereo fisheye image generation pipeline based on high-resolution LiDAR scans. Leveraging these methods, we thoroughly evaluate state-of-the-art monocular depth, stereo matching, and depth completion models on our benchmark. Additionally, we provide 18K LiDAR-derived sparse depth training samples, achieving up to a 62% performance boost on fisheye data when fine-tuning pinhole-based stereo models. In summary, the high precision and versatility of our benchmark set a strong foundation for advancing research in fisheye depth estimation and robotics perception. Project page: https://ilyaind.github.io/WideDepth
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。