arXiv:2411.17761cs.CV2024-11NeurIPS被引 19

首个面向3D目标检测的开放世界自动驾驶评测基准,支持复杂场景与未知物体识别。

OpenAD: Open-World Autonomous Driving Benchmark for 3D Object Detection

  • 基于多模态大模型构建异常场景发现与标注流程,覆盖5个数据集2000个场景。
  • 首次在真实开放世界设置下评估多种2D/3D模型,揭示现有方法在未知物体上的性能短板。
  • 提出视觉主导的基线模型并融合通用与专用模型,提升对罕见目标的检测精度。

开放世界感知旨在开发能适应新领域、多样化传感器配置,并理解罕见物体与极端情况的模型。然而,当前研究缺乏足够全面的开放世界3D感知评测基准和鲁棒的泛化方法。本文提出OpenAD,首个面向3D目标检测的真实开放世界自动驾驶评测基准。OpenAD基于集成多模态大语言模型(MLLM)的异常场景发现与标注流程,以统一格式标注了涵盖5个自动驾驶感知数据集的2000个场景。我们设计了评估方法,对多种开放世界及专用2D/3D模型进行了评估。此外,提出一种以视觉为中心的3D开放世界目标检测基线,并引入一种融合通用与专用模型的集成方法,以解决现有开放世界方法在该基准上精度偏低的问题。相关数据、工具代码与评估代码已开源,挑战赛已在EvalAI上线。

原文摘要 · Abstract (English)

Open-world perception aims to develop a model adaptable to novel domains and various sensor configurations and can understand uncommon objects and corner cases. However, current research lacks sufficiently comprehensive open-world 3D perception benchmarks and robust generalizable methodologies. This paper introduces OpenAD, the first real open-world autonomous driving benchmark for 3D object detection. OpenAD is built upon a corner case discovery and annotation pipeline that integrates with a multimodal large language model (MLLM). The proposed pipeline annotates corner case objects in a unified format for five autonomous driving perception datasets with 2000 scenarios. In addition, we devise evaluation methodologies and evaluate various open-world and specialized 2D and 3D models. Moreover, we propose a vision-centric 3D open-world object detection baseline and further introduce an ensemble method by fusing general and specialized models to address the issue of lower precision in existing open-world methods for the OpenAD benchmark. We host an online challenge on EvalAI. Data, toolkit codes, and evaluation codes are available at https://github.com/VDIGPKU/OpenAD.

3D检测开放世界自动驾驶多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。