arXiv:2411.05311cs.CVcs.RO2024-11NeurIPS被引 7

零样本跨模态感知框架,让自动驾驶场景自动标注更智能。

ZOPP: A Framework of Zero-shot Offboard Panoptic Perception for Autonomous Driving

  • 融合视觉大模型与点云3D表示,实现零样本识别
  • 在Waymo数据集上覆盖多种感知任务,效果优于传统方法
  • 适合需要快速适配新类别或少标注场景的自动驾驶研究

离线感知旨在为自动驾驶场景自动生成高质量3D标签。现有离线方法聚焦于封闭集分类的3D目标检测,难以匹配人类在快速演进的感知任务中的识别能力。由于高度依赖人工标注以及数据不平衡和稀疏性问题,尚未形成统一框架来满足各类感知任务对自动标注的需求。本文提出一种新型多模态零样本离线全景感知(ZOPP)框架,集成视觉基础模型的强大零样本识别能力与点云生成的3D表征。据我们所知,ZOPP是首个在多模态全景感知与自动驾驶场景自动标注领域的开创性工作。我们在Waymo Open Dataset上进行全面实验验证,涵盖多种感知任务,并进一步在下游应用中测试其可用性与可扩展性。结果充分展示了该框架在真实场景中的巨大潜力。

原文摘要 · Abstract (English)

Offboard perception aims to automatically generate high-quality 3D labels for autonomous driving (AD) scenes. Existing offboard methods focus on 3D object detection with closed-set taxonomy and fail to match human-level recognition capability on the rapidly evolving perception tasks. Due to heavy reliance on human labels and the prevalence of data imbalance and sparsity, a unified framework for offboard auto-labeling various elements in AD scenes that meets the distinct needs of perception tasks is not being fully explored. In this paper, we propose a novel multi-modal Zero-shot Offboard Panoptic Perception (ZOPP) framework for autonomous driving scenes. ZOPP integrates the powerful zero-shot recognition capabilities of vision foundation models and 3D representations derived from point clouds. To the best of our knowledge, ZOPP represents a pioneering effort in the domain of multi-modal panoptic perception and auto labeling for autonomous driving scenes. We conduct comprehensive empirical studies and evaluations on Waymo open dataset to validate the proposed ZOPP on various perception tasks. To further explore the usability and extensibility of our proposed ZOPP, we also conduct experiments in downstream applications. The results further demonstrate the great potential of our ZOPP for real-world scenarios.

自动驾驶零样本全景感知自标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。