首个融合视觉、触觉与3D点云的薄物体重建数据集,助力提升细长物体三维重建精度。
VisTa3D: A Dataset and Benchmark for Thin Object Reconstruction from Vision, Tactile, and 3D Point Clouds

- 构建多模态同步数据集,包含RGB、深度图和触觉响应图。
- 70个薄物体在17个场景中,共387个观测帧,含激光扫描真值。
- 首次验证触觉信息可显著改善细长物体的重建质量,适合机器人感知研究者。
当前先进的三维重建模型在处理薄物体时表现不佳,主要因这些物体在RGB图像和3D点云中占据空间极小。为评估其误差程度,我们构建了首个包含同步RGB图像、深度图和触觉响应图的数据集,每帧均附带惯性测量、相机位姿与标定信息,以及由激光扫描获得的真实深度与分割图。我们假设触觉数据能通过提供局部形状与形变信息来辅助薄物体重建。该数据集命名为VisTa3D,涵盖70个薄物体、17个环境中的387个场景。我们在该数据集上基准测试了现有3D重建模型,发现它们对薄物体的重建保真度普遍较低。为验证触觉数据的有效性,我们引入首个视觉-深度-触觉联合的三维重建模型作为基线。代码与数据:https://huggingface.co/datasets/shaniaguo/VisTa3D。
原文摘要 · Abstract (English)
State-of-the-art 3D reconstruction models, whether from visual, range, or both, tend to underperform on thin objects. This is partially due to the small amount of space such objects occupy in RGB images and in 3D point clouds. To test the extent of their errors, we collected the first thin object dataset comprising of synchronized RGB images, depth maps, and tactile response maps, where each frame is associated with inertial measurements, camera pose and calibration, and groundtruth depth and segmentation maps obtained from laser scanning of thin objects. We hypothesize that tactile data can aid in the reconstruction of thin objects as their response maps provide local shape and deformation information. Our dataset, termed VisTa3D, comprises of 387 scenes covering 70 thin objects over 17 environments. We benchmarked current 3D reconstruction models on VisTa3D and found that, indeed, they exhibit low fidelity on thin objects. To test if tactile data can help, we introduce the first visual-range-tactile 3D reconstruction model as a baseline. Code and data: https://huggingface.co/datasets/shaniaguo/VisTa3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。