arXiv:2412.13569cs.CV2024-12AAAI被引 9

构建新数据集与模型,实现多视角行人占用预测

Multi-View Pedestrian Occupancy Prediction with a Novel Synthetic Dataset

论文配图:Multi-View Pedestrian Occupancy Prediction with a Novel Synthetic Dataset
图 1 · 摘自论文原文
  • 用体素结构构建大规模合成数据集MVP-Occ
  • 提出OmniOcc模型,可预测全场景行人占据与语义标签
  • 适合城市交通感知与自动驾驶场景研究者

我们解决城市交通中多视角行人检测的进阶挑战——行人占用预测。为此,我们构建了名为MVP-Occ的新合成数据集,专为大场景密集行人场景设计。该数据集采用体素结构精细表征行人,并附带丰富的语义场景理解标签,有助于视觉导航与行人空间信息分析。此外,我们提出一个稳健的基线模型OmniOcc,可从多视角图像中预测整个场景的体素占据状态与全景语义标签。通过深入分析,我们识别并评估了模型的关键组件,明确了其具体贡献与重要性。

原文摘要 · Abstract (English)

We address an advanced challenge of predicting pedestrian occupancy as an extension of multi-view pedestrian detection in urban traffic. To support this, we have created a new synthetic dataset called MVP-Occ, designed for dense pedestrian scenarios in large-scale scenes. Our dataset provides detailed representations of pedestrians using voxel structures, accompanied by rich semantic scene understanding labels, facilitating visual navigation and insights into pedestrian spatial information. Furthermore, we present a robust baseline model, termed OmniOcc, capable of predicting both the voxel occupancy state and panoptic labels for the entire scene from multi-view images. Through in-depth analysis, we identify and evaluate the key elements of our proposed model, highlighting their specific contributions and importance.

行人预测多视角合成数据体素建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。