无需人工标注,实现多视角行人建模的精准深度一致性
DCHM: Depth-Consistent Human Modeling for Multiview Detection
- 基于超像素级高斯点云融合,保证多视角深度一致
- 在稀疏视图大场景下生成低噪点云,提升定位精度
- 首次在复杂场景中完成行人重建与多视角分割
多视角行人检测通常分为人体建模与行人定位两阶段。人体建模通过融合多视角信息在三维空间表示行人,其质量直接影响检测准确率。然而,现有方法常引入噪声且精度不足。尽管部分方法通过依赖昂贵的多视角3D标注降低噪声,却难以在多样场景中泛化。为此,我们提出深度一致人体建模(DCHM),一种无需人工标注、可实现全局坐标系下一致深度估计与多视角融合的框架。具体地,该方法采用超像素级高斯点云投射,在稀疏视角、大尺度及拥挤场景中实现多视角深度一致性,生成精确点云用于行人定位。大量实验表明,该方法显著降低了建模过程中的噪声,优于先前最先进方法。此外,据我们所知,DCHM是首个在如此挑战性环境下实现行人重建与多视角分割的工作。代码已开源。
原文摘要 · Abstract (English)
Multiview pedestrian detection typically involves two stages: human modeling and pedestrian localization. Human modeling represents pedestrians in 3D space by fusing multiview information, making its quality crucial for detection accuracy. However, existing methods often introduce noise and have low precision. While some approaches reduce noise by fitting on costly multiview 3D annotations, they often struggle to generalize across diverse scenes. To eliminate reliance on human-labeled annotations and accurately model humans, we propose Depth-Consistent Human Modeling (DCHM), a framework designed for consistent depth estimation and multiview fusion in global coordinates. Specifically, our proposed pipeline with superpixel-wise Gaussian Splatting achieves multiview depth consistency in sparse-view, large-scaled, and crowded scenarios, producing precise point clouds for pedestrian localization. Extensive validations demonstrate that our method significantly reduces noise during human modeling, outperforming previous state-of-the-art baselines. Additionally, to our knowledge, DCHM is the first to reconstruct pedestrians and perform multiview segmentation in such a challenging setting. Code is available on the \href{https://jiahao-ma.github.io/DCHM/}{project page}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。