构建首个面向半结构化场景的多模态行人数据集,提升复杂环境感知能力。
PFSD: A Multi-Modal Pedestrian-Focus Scene Dataset for Rich Tasks in Semi-Structured Environments
- 构建半结构化场景下的多模态行人数据集,含13万+行人实例与精细标注
- 提出混合多尺度融合网络,3D行人检测mAP显著优于现有方法
- 适合自动驾驶感知、行人预测与复杂场景建模研究者使用
近年来自动驾驶感知在以车辆为主导的结构化环境中表现出色,但在动态行人多样、遮挡频繁的半结构化环境中仍存在明显局限。我们归因于此类场景高质量数据集的匮乏,尤其是针对行人感知与预测的数据。本文提出多模态行人聚焦场景数据集PFSD,基于nuScenes格式,在半结构化场景中进行严格标注,涵盖超过13万例行人实例,覆盖不同密度、运动模式与遮挡情况。数据集提供点云分割、检测及物体追踪ID等多模态标注。为应对复杂半结构化环境挑战,我们提出一种新型混合多尺度融合网络(HMFN),通过结合稀疏卷积与常规卷积的混合框架,有效捕捉并融合多尺度特征。在PFSD上的大量实验表明,该方法在3D行人检测任务上实现了比现有方法更高的均值平均精度(mAP),验证了其在复杂场景中的有效性。代码与基准测试已公开。
原文摘要 · Abstract (English)
Recent advancements in autonomous driving perception have revealed exceptional capabilities within structured environments dominated by vehicular traffic. However, current perception models exhibit significant limitations in semi-structured environments, where dynamic pedestrians with more diverse irregular movement and occlusion prevail. We attribute this shortcoming to the scarcity of high-quality datasets in semi-structured scenes, particularly concerning pedestrian perception and prediction. In this work, we present the multi-modal Pedestrian-Focused Scene Dataset(PFSD), rigorously annotated in semi-structured scenes with the format of nuScenes. PFSD provides comprehensive multi-modal data annotations with point cloud segmentation, detection, and object IDs for tracking. It encompasses over 130,000 pedestrian instances captured across various scenarios with varying densities, movement patterns, and occlusions. Furthermore, to demonstrate the importance of addressing the challenges posed by more diverse and complex semi-structured environments, we propose a novel Hybrid Multi-Scale Fusion Network (HMFN). Specifically, to detect pedestrians in densely populated and occluded scenarios, our method effectively captures and fuses multi-scale features using a meticulously designed hybrid framework that integrates sparse and vanilla convolutions. Extensive experiments on PFSD demonstrate that HMFN attains improvement in mean Average Precision (mAP) over existing methods, thereby underscoring its efficacy in addressing the challenges of 3D pedestrian detection in complex semi-structured environments. Coding and benchmark are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。