用视觉语言模型自动推断车辆信息,提升自动驾驶3D标注精度与效率。
Improving 3D Labeling in Self-Driving by Inferring Vehicle Information using Vision Language Models

- 利用视觉语言模型从图像中推断车型、品牌和年份,生成精准3D框尺寸。
- 在严重遮挡场景下,模型推算的尺寸优于激光雷达辅助的人工标注。
- 可显著减少人工标注时间,适用于不同数据集和标注者,通用性强。
我们提出一种通过零样本推理车辆信息来改进自动驾驶中3D车辆标注的方法,利用车辆品牌与型号识别(VMMR)技术。该方法采用视觉语言模型(VLM)从图像裁片中同时推断车辆的品牌、型号及年代,并输出准确的3D边界框尺寸以辅助人工标注。我们评估了迭代提示工程和不同VLM选择对车辆边界框推断及品牌/型号/年份识别的影响。与强基线相比,该方法不仅精度高,还能有效缓解特定失效模式——在车辆严重遮挡情况下,VLM推算的尺寸优于初始激光雷达辅助的人工标注。在公开和私有数据上的实验均表明,结论具有跨标注者和数据集的泛化能力。结果表明,将VLM融入标注流程可降低人工标注时间并提升标注质量。
原文摘要 · Abstract (English)
We present an approach to improve 3D vehicle labeling in self-driving applications through zero-shot inference of vehicle information, leveraging Vehicle Make and Model Recognition (VMMR) methods. The proposed approach utilizes a Vision Language Model (VLM) to both infer a vehicle's make, model, and generation from image crops, and output accurate 3D bounding box dimensions to seed manual labeling. We evaluate the impact of iterative prompt engineering and the choice of different VLMs on both vehicle bounding box inference and make/model/generation recognition. When compared to strong baselines, the proposed approach not only shows high accuracy, but also excels in mitigating specific failure modes where VLMs provide better dimensions than initial lidar-aided human annotated labels (e.g., in cases of significant vehicle occlusion). Experiments on both public and proprietary data strongly suggest that our conclusions are generalizable across different labelers and datasets. The results demonstrate that integrating VLMs into the labeling process can reduce manual labeling time while increasing label quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。