arXiv:2410.01678cs.CVcs.RO2024-10ICRA被引 5

让自动驾驶系统能追踪没见过的物体,突破传统识别类别限制。

Open3DTrack: Towards Open-Vocabulary 3D Multi-Object Tracking

  • 将开放词汇能力融入3D跟踪框架,支持未见物体识别。
  • 在多种真实驾驶场景中实现对新类物体的有效追踪。
  • 首个面向开放词汇3D多目标追踪的工作,适合自动驾驶研究者。

3D多目标跟踪在自动驾驶中至关重要,可实现实时监控与预测多个物体的运动。传统3D跟踪系统受限于预定义物体类别,难以适应动态环境中的新出现物体。为此,本文提出开放词汇3D跟踪,扩展跟踪范围至预定义类别之外的物体。我们定义了开放词汇3D跟踪问题,并设计了涵盖不同开放场景的数据集划分。提出一种新方法,将开放词汇能力集成到3D跟踪框架中,实现对未见物体类别的泛化。通过策略性适配,有效缩小已知与新物体之间的性能差距。实验表明,该方法在多样化的室外驾驶场景中具备鲁棒性与适应性。据我们所知,这是首个解决开放词汇3D跟踪的工作,为真实场景下的自动驾驶系统带来重要进展。代码、训练模型及数据集划分均已公开。

原文摘要 · Abstract (English)

3D multi-object tracking plays a critical role in autonomous driving by enabling the real-time monitoring and prediction of multiple objects' movements. Traditional 3D tracking systems are typically constrained by predefined object categories, limiting their adaptability to novel, unseen objects in dynamic environments. To address this limitation, we introduce open-vocabulary 3D tracking, which extends the scope of 3D tracking to include objects beyond predefined categories. We formulate the problem of open-vocabulary 3D tracking and introduce dataset splits designed to represent various open-vocabulary scenarios. We propose a novel approach that integrates open-vocabulary capabilities into a 3D tracking framework, allowing for generalization to unseen object classes. Our method effectively reduces the performance gap between tracking known and novel objects through strategic adaptation. Experimental results demonstrate the robustness and adaptability of our method in diverse outdoor driving scenarios. To the best of our knowledge, this work is the first to address open-vocabulary 3D tracking, presenting a significant advancement for autonomous systems in real-world settings. Code, trained models, and dataset splits are available publicly.

3D跟踪开放词汇自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。