修复内窥镜深度数据集中的相机位姿错误,大幅扩充可用数据量
SCARED-C: Corrected Camera Poses for Endoscopic Depth Estimation

- 用COLMAP重估所有帧的相机位姿,结合真值深度图恢复尺度
- 可靠RGB-D样本从35对增至17,135对,提升近500倍
- 适用于内窥镜深度估计、视觉里程计等任务的研究者
SCARED数据集是内窥镜深度估计的常用基准,提供由结构光传感器捕获的真值3D重建。然而,非关键帧的深度图依赖机器人运动学估算,引入显著位姿误差,导致仅35个关键帧可信赖。本文提出SCARED-C,通过COLMAP系统重新估计所有帧的相机位姿,并利用真值关键帧深度图进行尺度恢复,将可靠RGB-D对数量从35对扩展至17,135对。通过立体视差评估与单目深度估计实验验证了修正后位姿的有效性。相关数据集与代码已公开发布。
原文摘要 · Abstract (English)
The SCARED dataset is a widely used benchmark for endoscopic depth estimation, offering ground-truth 3D reconstructions captured with a structured light sensor. However, the depth maps for non-keyframe images rely on robot kinematics that introduce substantial pose errors, limiting the reliably labeled portion of the dataset to 35 keyframes. We present SCARED-C, a corrected version of the SCARED dataset that expands the number of reliable RGB-D pairs from 35 to 17,135. Our pipeline applies COLMAP, a Structure-from-Motion system, to re-estimate camera poses for all frames, followed by a scale recovery step that aligns the resulting reconstructions to metric space using the ground-truth keyframe depth maps. We validate the corrected poses through (1) stereo disparity evaluation and (2) monocular depth estimation experiments. The corrected dataset and code are publicly released to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。