arXiv:2504.08154cs.CV2025-04被引 3

用视觉语言模型分析激光点云,实现少标注下的卡车精准分类。

Investigating Vision-Language Model for Point Cloud-based Vehicle Classification

  • 将点云数据通过注册与形态学处理,适配视觉语言模型输入
  • 仅需少量标注样本,即可在真实路侧点云上实现高精度分类
  • 适合自动驾驶中需快速部署的车辆识别场景

重型卡车因体积大、机动性差,对协同自动驾驶安全构成重大挑战。传统基于激光雷达的卡车分类依赖大量人工标注,成本高昂。大型语言模型具备少样本学习能力,但现有视觉语言模型主要基于图像训练,难以直接处理点云数据。本研究提出一种新框架,融合路边激光雷达点云与视觉语言模型,实现高效准确的卡车分类,支持协同安全驾驶。创新包括:(1)使用真实世界激光雷达数据集进行模型开发;(2)设计预处理流程,通过点云配准实现密集3D渲染,并采用数学形态学增强特征表达;(3)利用上下文学习与少样本提示,实现最小标注下的车辆分类。实验表明该方法性能优异,有望显著降低标注成本并提升分类精度。

原文摘要 · Abstract (English)

Heavy-duty trucks pose significant safety challenges due to their large size and limited maneuverability compared to passenger vehicles. A deeper understanding of truck characteristics is essential for enhancing the safety perspective of cooperative autonomous driving. Traditional LiDAR-based truck classification methods rely on extensive manual annotations, which makes them labor-intensive and costly. The rapid advancement of large language models (LLMs) trained on massive datasets presents an opportunity to leverage their few-shot learning capabilities for truck classification. However, existing vision-language models (VLMs) are primarily trained on image datasets, which makes it challenging to directly process point cloud data. This study introduces a novel framework that integrates roadside LiDAR point cloud data with VLMs to facilitate efficient and accurate truck classification, which supports cooperative and safe driving environments. This study introduces three key innovations: (1) leveraging real-world LiDAR datasets for model development, (2) designing a preprocessing pipeline to adapt point cloud data for VLM input, including point cloud registration for dense 3D rendering and mathematical morphological techniques to enhance feature representation, and (3) utilizing in-context learning with few-shot prompting to enable vehicle classification with minimally labeled training data. Experimental results demonstrate encouraging performance of this method and present its potential to reduce annotation efforts while improving classification accuracy.

点云分类视觉语言模型少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。