arXiv:2411.18025cs.CV2024-11CVPR被引 12

提出像素对齐的双目可见光与近红外成像系统及数据集

Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision

  • 构建移动机器人平台实现可见光与近红外图像像素级对齐
  • 在多种光照条件下验证了融合方法在3D视觉中的有效性
  • 适合做机器人感知、多模态视觉融合的研究者参考

将可见光(RGB)与近红外(NIR)双目成像结合可提供互补的光谱信息,有望提升机器人在复杂光照条件下的三维视觉能力。然而现有数据集与成像系统缺乏RGB与NIR图像间的像素级对齐,影响下游视觉任务性能。本文提出一种搭载像素对齐的RGB-NIR双目相机与LiDAR传感器的移动机器人视觉系统,可同步采集像素对齐的左右视图RGB与NIR图像以及时间对齐的LiDAR点云。借助机器人移动性,构建了覆盖多种光照条件的连续视频数据集。进一步提出两种利用像素对齐的RGB-NIR图像的方法:一是无需微调即可让预训练的RGB模型直接使用多模态信息;二是通过微调使模型更高效地融合多模态特征。实验表明,在多样化光照条件下,该方法显著提升了视觉性能。

原文摘要 · Abstract (English)

Integrating RGB and NIR stereo imaging provides complementary spectral information, potentially enhancing robotic 3D vision in challenging lighting conditions. However, existing datasets and imaging systems lack pixel-level alignment between RGB and NIR images, posing challenges for downstream vision tasks. In this paper, we introduce a robotic vision system equipped with pixel-aligned RGB-NIR stereo cameras and a LiDAR sensor mounted on a mobile robot. The system simultaneously captures pixel-aligned pairs of RGB stereo images, NIR stereo images, and temporally synchronized LiDAR points. Utilizing the mobility of the robot, we present a dataset containing continuous video frames under diverse lighting conditions. We then introduce two methods that utilize the pixel-aligned RGB-NIR images: an RGB-NIR image fusion method and a feature fusion method. The first approach enables existing RGB-pretrained vision models to directly utilize RGB-NIR information without fine-tuning. The second approach fine-tunes existing vision models to more effectively utilize RGB-NIR information. Experimental results demonstrate the effectiveness of using pixel-aligned RGB-NIR images across diverse lighting conditions.

多模态融合机器人视觉立体成像数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。