用概率占位体积提升多视角行人检测精度
Enhanced Multi-View Pedestrian Detection Using Probabilistic Occupancy Volume
- 引入概率占位体积聚焦行人区域,优化3D特征构建
- 在MultiviewX上达97.3% MODA,超越现有模型
- 适合多摄像头行人检测场景,尤其遮挡严重时
单视角行人检测受遮挡影响严重。多视角系统通过融合多视角信息缓解此问题。现有方法采用早期融合策略,将特征投影至地面平面进行检测分析。一种有前景的方法是3D特征拉取技术,通过采样每个体素对应的2D特征构建整个场景的3D特征体积。然而,该方法未考虑行人可能存在的位置,导致特征冗余。本文提出新模型,结合传统3D重建技术,通过引入基于视觉外壳技术构建的概率占位体积,补充3D特征体积。该体积聚焦于可能含行人的区域,引导模型注意力,提升检测精度。在MultiviewX数据集上,模型达到97.3% MODA,优于当前最优模型;在Wildtrack数据集上表现相当。
原文摘要 · Abstract (English)
Occlusion poses a significant challenge in pedestrian detection from a single view. To address this, multi-view detection systems have been utilized to aggregate information from multiple perspectives. Recent advances in multi-view detection utilized an early-fusion strategy that strategically projects the features onto the ground plane, where detection analysis is performed. A promising approach in this context is the use of 3D feature-pulling technique, which constructs a 3D feature volume of the scene by sampling the corresponding 2D features for each voxel. However, it creates a 3D feature volume of the whole scene without considering the potential locations of pedestrians. In this paper, we introduce a novel model that efficiently leverages traditional 3D reconstruction techniques to enhance deep multi-view pedestrian detection. This is accomplished by complementing the 3D feature volume with probabilistic occupancy volume, which is constructed using the visual hull technique. The probabilistic occupancy volume focuses the model's attention on regions occupied by pedestrians and improves detection accuracy. Our model outperforms state-of-the-art models on the MultiviewX dataset, with an MODA of 97.3%, while achieving competitive performance on the Wildtrack dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。