arXiv:2411.06232cs.CV2024-11被引 2

从单张大场景图中精准重建数百人姿态与体型,且不受视角和尺度影响。

RCR: Robust Crowd Reconstruction with Upright Space from a Single Large-scene Image

  • 用虚拟交互点将3D定位转为2D像素定位,解决小尺度人体难题。
  • 在真实与合成数据集上实现全局一致的重建,无需测试时优化。
  • 适合大规模人群重建、自动驾驶与监控场景的视觉理解任务。

本文致力于从单张大场景图像中,实现数百人的空间一致性姿态与体型重建,适用于任意相机视场角(FoV)及多种人体尺度。由于人体在图像中尺寸小且变化大、存在深度模糊与透视畸变,现有方法无法实现全局一致且正确重投影的重建。为此,我们提出新概念“人类-场景虚拟交互点”(HVIP),将复杂的3D人体定位转化为2D像素定位。在此基础上构建鲁棒人群重建框架RCR,可在不同相机视场角下实现全局一致重建并稳定泛化,无需测试时优化。为应对不同像素尺度的人体感知,提出迭代地面感知裁剪策略,自动裁剪图像并融合结果。为消除相机与裁剪过程的影响,引入规范直立3D空间与对应直立2D空间,并设计直立归一化机制,将局部裁剪输入映射到直立2D空间,再将输出从直立3D空间转换回统一相机空间。此外,我们构建了两个基准数据集LargeCrowd和SynCrowd,用于评估大场景下人群重建性能。实验验证了方法的有效性,源代码与数据将公开供研究使用。

原文摘要 · Abstract (English)

This paper focuses on spatially consistent hundreds of human pose and shape reconstruction from a single large-scene image with various human scales under arbitrary camera FoVs (Fields of View). Due to the small and highly varying 2D human scales, depth ambiguity, and perspective distortion, no existing methods can achieve globally consistent reconstruction with correct reprojection. To address these challenges, we first propose a new concept, Human-scene Virtual Interaction Point (HVIP), to convert the complex 3D human localization into 2D-pixel localization. We then extend it to RCR (Robust Crowd Reconstruction), which achieves globally consistent reconstruction and stable generalization on different camera FoVs without test-time optimization. To perceive humans in varying pixel sizes, we propose an Iterative Ground-aware Cropping to automatically crop the image and then merge the results. To eliminate the influence of the camera and cropping process during the reconstruction, we introduce a canonical Upright 3D Space and the corresponding Upright 2D Space. To link the canonical space and the camera space, we propose the Upright Normalization, which transforms the local crop input into the Upright 2D Space, and transforms the output from the Upright 3D Space into the unified camera space. Besides, we contribute two benchmark datasets, LargeCrowd and SynCrowd, for evaluating crowd reconstruction in large scenes. Experimental results demonstrate the effectiveness of the proposed method. The source code and data will be publicly available for research purposes.

人群重建单图像姿态估计3D建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。