针对密集人群场景,提升激光雷达-相机3D跟踪的表征能力与标注效率。
Learning better representations for crowded pedestrians in offboard LiDAR-camera 3D tracking-by-detection
- 设计密度与关系感知的高分辨率表征,增强小目标和拥挤场景下的建模能力。
- 在新构建的多视角密集行人数据集上,3D跟踪性能显著提升,标注效率更高。
- 适用于自动驾驶中复杂城市环境的行人感知系统优化,尤其适合研究密集场景建模。
基于学习的自主感知在高度拥挤的城市环境中感知行人仍属长期尾部难题。加速此类挑战性场景的3D真实标注生成对性能至关重要但极富挑战。主要难点包括行人点云稀疏以及缺乏适用于特定系统设计研究的基准数据集。为此,我们首先收集了一个新的、面向高度拥挤行人的多视角激光雷达-相机3D多目标跟踪基准数据集,用于深入分析。随后构建了一个离线自动标注系统,通过激光雷达点云与多视角图像重建行人轨迹。为提升密集场景下的泛化能力及小物体检测性能,我们提出学习具有密度感知与关系感知特性的高分辨率表征。大量实验验证了该方法能显著提升3D行人跟踪性能,并实现更高的自动标注效率。代码将公开于该网址。
原文摘要 · Abstract (English)
Perceiving pedestrians in highly crowded urban environments is a difficult long-tail problem for learning-based autonomous perception. Speeding up 3D ground truth generation for such challenging scenes is performance-critical yet very challenging. The difficulties include the sparsity of the captured pedestrian point cloud and a lack of suitable benchmarks for a specific system design study. To tackle the challenges, we first collect a new multi-view LiDAR-camera 3D multiple-object-tracking benchmark of highly crowded pedestrians for in-depth analysis. We then build an offboard auto-labeling system that reconstructs pedestrian trajectories from LiDAR point cloud and multi-view images. To improve the generalization power for crowded scenes and the performance for small objects, we propose to learn high-resolution representations that are density-aware and relationship-aware. Extensive experiments validate that our approach significantly improves the 3D pedestrian tracking performance towards higher auto-labeling efficiency. The code will be publicly available at this HTTP URL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。