arXiv:2601.01222cs.CV2026-01被引 7

一次前向计算完成高精度人与场景三维重建,解决真实数据少的问题。

UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass

  • 融合场景与人体先验,用无标签真实视频训练提升泛化能力。
  • 通过深度模型蒸馏细节,使人体表面更精细,几何对齐更准确。
  • 适合做真实世界中人物主导的3D重建,如影视、虚拟人应用。

我们提出UniSH,一个统一的前向传播框架,用于联合进行度量尺度下的三维场景与人体重建。该领域的主要挑战是缺乏大规模标注的真实世界数据,导致依赖合成数据,引入显著的模拟到现实域差距,造成泛化能力差、人体几何质量低、真实视频中对齐效果不佳。为解决此问题,我们设计了一种创新训练范式,有效利用未标注的真实世界数据。框架融合了场景重建与人体运动恢复(HMR)中的强先验,并通过两个核心组件训练:(1) 一种鲁棒的蒸馏策略,从专家深度模型中提取高频细节以细化人体表面;(2) 两阶段监督机制,先在合成数据上学习粗略定位,再在真实数据上通过直接优化SMPL网格与人体点云间的几何对应关系进行微调。该方法使前向模型能在单次前向传播中同时恢复高保真场景几何、人体点云、相机参数以及一致的度量尺度下SMPL人体模型。大量实验表明,本模型在以人为中心的场景重建任务上达到最先进性能,并在全局人体动作估计上表现优异,优于基于优化的方法和仅人体重建方法。

原文摘要 · Abstract (English)

We present UniSH, a unified, feed-forward framework for joint metric-scale 3D scene and human reconstruction. A key challenge in this domain is the scarcity of large-scale, annotated real-world data, forcing a reliance on synthetic datasets. This reliance introduces a significant sim-to-real domain gap, leading to poor generalization, low-fidelity human geometry, and poor alignment on in-the-wild videos. To address this, we propose an innovative training paradigm that effectively leverages unlabeled in-the-wild data. Our framework bridges strong, disparate priors from scene reconstruction and HMR, and is trained with two core components: (1) a robust distillation strategy to refine human surface details by distilling high-frequency details from an expert depth model, and (2) a two-stage supervision scheme, which first learns coarse localization on synthetic data, then fine-tunes on real data by directly optimizing the geometric correspondence between the SMPL mesh and the human point cloud. This approach enables our feed-forward model to jointly recover high-fidelity scene geometry, human point clouds, camera parameters, and coherent, metric-scale SMPL bodies, all in a single forward pass. Extensive experiments demonstrate that our model achieves state-of-the-art performance on human-centric scene reconstruction and delivers highly competitive results on global human motion estimation, comparing favorably against both optimization-based frameworks and HMR-only methods. Project page: https://murphylmf.github.io/UniSH/

三维重建人体建模前向传播真实数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。