arXiv:2607.19517cs.CV2026-07

从单目视频重建大规模场景中一致的4D人群运动,突破深度模糊限制。

Crowd4D: Scene-Aware Monocular 4D Crowd Reconstruction

论文配图:Crowd4D: Scene-Aware Monocular 4D Crowd Reconstruction
图 1 · 摘自论文原文
  • 通过人体-场景交互代理联合优化人群与场景几何
  • 在复杂地形下实现无漂移、高精度的尺度一致性重建
  • 适合需要真实尺度人群动态建模的智能交通与安防应用

从单目视频中恢复大规模场景下的一致性4D人群运动仍面临严重深度模糊和复杂场景几何的挑战。现有方法通常依赖单平面假设,导致在复杂地形下度量尺度不可靠且空间漂移。我们提出Crowd4D,首个面向场景感知的4D人群重建框架,可从单目RGB视频中联合优化人群与场景。该方法通过多阶段优化策略显式融合场景几何,确保图像空间与场景空间的一致性。关键瓶颈在于准确的人体-场景对齐(尺度与位置),但传统方法将两者解耦。为此,我们引入人体-场景交互代理(HSIP),基于场景交互点云(SIPC)和场景交互表面(SIS)构建中间表示,编码显式的场景感知几何先验,并重构大尺度单目4D人群重建的优化空间。为提升遮挡下的时间稳定性,我们进一步提出人群结构一致性正则化(CSCR),利用基于HSIP的空间先验,对局部人群邻域内的相对位移与方向施加软性时序一致性约束。大量实验表明,Crowd4D持续优于现有最先进方法,在复杂真实场景中实现鲁棒的单目4D人群重建。

原文摘要 · Abstract (English)

Recovering scene-consistent 4D crowd motion from monocular video in large-scale scenes remains challenging due to severe depth ambiguity and complex scene geometry. Existing monocular crowd reconstruction methods typically rely on single-plane assumptions, leading to unreliable metric scale and spatial drift under complex terrain. We propose Crowd4D, the first scene-aware 4D crowd reconstruction framework that jointly optimizes the crowd and scene from a monocular RGB video in large-scale scenes. Crowd4D explicitly incorporates scene geometry and ensures consistency across image and scene spaces via a multi-stage optimization strategy. A key bottleneck of this task lies in accurate human-scene alignment, particularly in scale and position. However, human and scene reconstructions are typically decoupled. To address this, we introduce the Human-Scene Interaction Proxy, abbreviated as HSIP, as an intermediate representation derived from Scene Interaction Point Clouds and a Scene Interaction Surface, abbreviated as SIPC and SIS. These representations encode explicit scene-aware geometric priors and redefine the optimization space for large-scale monocular 4D crowd reconstruction. To further improve temporal stability under occlusions, we introduce Crowd Structural Coherence Regularization, abbreviated as CSCR, which leverages HSIP-based spatial priors to impose soft temporal consistency on pairwise relative displacements and directions within local crowd neighborhoods. Extensive experiments demonstrate that Crowd4D consistently outperforms existing state-of-the-art methods and enables robust monocular 4D crowd reconstruction in complex, large-scale real-world scenes.

4D重建单目视觉人群建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。