arXiv:2508.12644cs.CV2025-08TPAMI被引 8

首个多视角视频动态人群3D重建框架,解决遮挡与时间不一致问题。

DyCrowd: Towards Dynamic Crowd Reconstruction from a Large-scene Video

  • 分阶段群体引导运动优化,利用集体行为缓解长期遮挡。
  • 引入变分自编码器人体运动先验与异步一致性损失,提升重建质量。
  • 适用于城市监控、人流分析等大规模场景,适合计算机视觉研究者。

大规模场景中动态人群的三维重建在城市监控与人群分析中日益重要。然而,现有方法多基于静态图像重建,导致时间不一致且难以缓解遮挡影响。本文提出 DyCrowd,首个实现数百人姿态、位置与形状在时空上一致的大型场景视频三维重建框架。设计了粗到精的群体引导运动优化策略,以应对大场景中的遮挡问题。进一步融合变分自编码器(VAE)人体运动先验与段级群体引导优化,核心思想是利用集体行为处理长期动态遮挡。通过联合优化相似运动片段中个体的运动序列,并结合提出的异步运动一致性(AMC)损失,使未被遮挡的高质量运动片段指导被遮挡部分的恢复,确保在时间不同步与节奏不一致情况下仍具鲁棒性与合理性。此外,为填补缺乏标注良好的大型场景视频数据集的空白,我们构建了虚拟基准数据集 VirtualCrowd,用于评估大规模动态人群重建。实验表明,该方法在大规模动态人群重建任务中达到领先性能。代码与数据集将开放供研究使用。

原文摘要 · Abstract (English)

3D reconstruction of dynamic crowds in large scenes has become increasingly important for applications such as city surveillance and crowd analysis. However, current works attempt to reconstruct 3D crowds from a static image, causing a lack of temporal consistency and inability to alleviate the typical impact caused by occlusions. In this paper, we propose DyCrowd, the first framework for spatio-temporally consistent 3D reconstruction of hundreds of individuals' poses, positions and shapes from a large-scene video. We design a coarse-to-fine group-guided motion optimization strategy for occlusion-robust crowd reconstruction in large scenes. To address temporal instability and severe occlusions, we further incorporate a VAE (Variational Autoencoder)-based human motion prior along with a segment-level group-guided optimization. The core of our strategy leverages collective crowd behavior to address long-term dynamic occlusions. By jointly optimizing the motion sequences of individuals with similar motion segments and combining this with the proposed Asynchronous Motion Consistency (AMC) loss, we enable high-quality unoccluded motion segments to guide the motion recovery of occluded ones, ensuring robust and plausible motion recovery even in the presence of temporal desynchronization and rhythmic inconsistencies. Additionally, in order to fill the gap of no existing well-annotated large-scene video dataset, we contribute a virtual benchmark dataset, VirtualCrowd, for evaluating dynamic crowd reconstruction from large-scene videos. Experimental results demonstrate that the proposed method achieves state-of-the-art performance in the large-scene dynamic crowd reconstruction task. The code and dataset will be available for research purposes.

人群重建三维重建视频理解运动估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。