arXiv:2510.06219cs.CV2025-10被引 35

单视频实时重建多人3D人体与场景,一步到位。

Human3R: Everyone Everywhere All at Once

  • 单次前向传播同步恢复多人3D身体、场景和相机轨迹。
  • 15帧/秒实时运行,仅需8GB显存,一天训练完成。
  • 无需检测、深度估计等预处理,适合快速部署应用。

我们提出Human3R,一种统一的前馈框架,可从随意拍摄的单目视频中在线重建世界坐标系下的4D人体-场景结构。不同于以往依赖多阶段流水线、迭代接触感知优化及重依赖(如人体检测、深度估计、SLAM预处理)的方法,Human3R在一次前向传递中联合恢复全局多人体SMPL-X模型(“每个人”)、稠密3D场景(“每个地方”)和相机轨迹(“同时”)。该方法基于4D在线重建模型CUT3R,采用参数高效的视觉提示调优,以保留CUT3R丰富的时空先验,同时实现多个SMPL-X人体的直接读取。经过仅一天、单张GPU在小规模合成数据集BEDLAM上的训练,其性能优越且效率极高:支持单次处理多人及3D场景重建,实时运行达15 FPS,内存占用低至8 GB。大量实验表明,Human3R在全局人体运动估计、局部人体网格恢复、视频深度估计和相机位姿估计等任务上均达到或超过当前最佳表现,且仅需单一统一模型。我们希望Human3R能成为简单而强大的基线,便于下游应用适配。代码、模型与4D交互演示已公开于https://fanegg.github.io/Human3R/。

原文摘要 · Abstract (English)

We present Human3R, a unified, feed-forward framework for online 4D human-scene reconstruction, in the world frame, from casually captured monocular videos. Unlike previous approaches that rely on multi-stage pipelines, iterative contact-aware refinement between humans and scenes, and heavy dependencies, e.g., human detection, depth estimation, and SLAM pre-processing, Human3R jointly recovers global multi-person SMPL-X bodies ("everyone"), dense 3D scene ("everywhere"), and camera trajectories in a single forward pass ("all-at-once"). Our method builds upon the 4D online reconstruction model CUT3R, and uses parameter-efficient visual prompt tuning, to strive to preserve CUT3R's rich spatiotemporal priors, while enabling direct readout of multiple SMPL-X bodies. Human3R is a unified model that eliminates heavy dependencies and iterative refinement. After being trained on the relatively small-scale synthetic dataset BEDLAM for just one day on one GPU, it achieves superior performance with remarkable efficiency: it reconstructs multiple humans in a one-shot manner, along with 3D scenes, in one stage, in real-time (15 FPS) with a low memory footprint (8 GB). Extensive experiments demonstrate that Human3R delivers state-of-the-art or competitive performance across tasks, including global human motion estimation, local human mesh recovery, video depth estimation, and camera pose estimation, with a single unified model. We hope that Human3R will serve as a simple yet strong baseline, which can be easily adapted for downstream applications. Code, models and 4D interactive demos are available at https://fanegg.github.io/Human3R/.

4D重建单目视频实时渲染多人建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。