arXiv:2502.04385cs.CVcs.AI2025-02被引 2

用激光雷达生成的2D图像实现文本理解,提升检测精度。

TexLiDAR: Automated Text Understanding for Panoramic LiDAR Data

  • 用OS1传感器的2D图像替代3D点云进行文本关联
  • Florence 2零样本生成更丰富描述,检测性能优于CLIP
  • 适合实时高精度场景,无需复杂点云处理

将激光雷达数据与文本关联的研究如LidarCLIP主要依赖3D点云嵌入到CLIP的图文空间,但点云编码效率低且难处理。随着Ouster OS1等先进传感器的出现,其不仅生成3D点云,还输出固定分辨率的深度、信号和环境全景2D图像,为基于激光雷达的任务带来新机遇。本文提出一种替代方案:利用OS1生成的2D图像而非3D点云连接文本。在零样本设置下,采用Florence 2大模型进行图像描述生成与目标检测。实验表明,Florence 2生成的描述更丰富,目标检测性能优于现有方法(如CLIP)。结合先进激光雷达数据与预训练大模型,该方法为高精度、强鲁棒性的挑战性检测场景(包括实时应用)提供了可靠解决方案。

原文摘要 · Abstract (English)

Efforts to connect LiDAR data with text, such as LidarCLIP, have primarily focused on embedding 3D point clouds into CLIP text-image space. However, these approaches rely on 3D point clouds, which present challenges in encoding efficiency and neural network processing. With the advent of advanced LiDAR sensors like Ouster OS1, which, in addition to 3D point clouds, produce fixed resolution depth, signal, and ambient panoramic 2D images, new opportunities emerge for LiDAR based tasks. In this work, we propose an alternative approach to connect LiDAR data with text by leveraging 2D imagery generated by the OS1 sensor instead of 3D point clouds. Using the Florence 2 large model in a zero-shot setting, we perform image captioning and object detection. Our experiments demonstrate that Florence 2 generates more informative captions and achieves superior performance in object detection tasks compared to existing methods like CLIP. By combining advanced LiDAR sensor data with a large pre-trained model, our approach provides a robust and accurate solution for challenging detection scenarios, including real-time applications requiring high accuracy and robustness.

激光雷达文本理解大模型目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。