arXiv:2603.12789cs.CV2026-03

单次通过实现多人多视角人与场景协同重建

Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass

  • 统一框架联合估计相机、场景点云和人体网格
  • 速度比以往方法快8倍以上,精度达领先水平
  • 适合需要实时多人3D重建的应用场景

近期3D基础模型的发展推动了对人体及其环境重建的兴趣。然而,多数现有方法依赖单目输入,扩展到多视角场景需额外模块或预处理数据。为此,我们提出CHROMM,一个统一框架,无需外部模块或预处理,即可从多人多视角视频中联合估计相机、场景点云和人体网格。该框架整合了Pi3X和Multi-HMR中的强几何与人体先验,构建可训练的神经网络,并引入尺度调整模块解决人体与场景间的尺度差异。同时设计多视角融合策略,在测试时将各视角估计聚合为统一表示。最后提出基于几何的多人关联方法,比外观方法更鲁棒。在EMDB、RICH、EgoHumans和EgoExo4D数据集上的实验表明,CHROMM在全局人体运动和多视角姿态估计上表现优异,且运行速度超过先前优化类多视角方法8倍以上。

原文摘要 · Abstract (English)

Recent advances in 3D foundation models have led to growing interest in reconstructing humans and their surrounding environments. However, most existing approaches focus on monocular inputs, and extending them to multi-view settings requires additional overhead modules or preprocessed data. To this end, we present CHROMM, a unified framework that jointly estimates cameras, scene point clouds, and human meshes from multi-person multi-view videos without relying on external modules or preprocessing. We integrate strong geometric and human priors from Pi3X and Multi-HMR into a single trainable neural network architecture, and introduce a scale adjustment module to solve the scale discrepancy between humans and the scene. We also introduce a multi-view fusion strategy to aggregate per-view estimates into a single representation at test-time. Finally, we propose a geometry-based multi-person association method, which is more robust than appearance-based approaches. Experiments on EMDB, RICH, EgoHumans, and EgoExo4D show that CHROMM achieves competitive performance in global human motion and multi-view pose estimation while running over 8x faster than prior optimization-based multi-view approaches. Project page: https://nstar1125.github.io/chromm.

3D重建多视角人体建模实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。