通过显式建模人体与场景接触,实现单目视频下高精度4D人体-场景重建。
UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video

- 基于人体姿态与场景几何推断4D接触关系,作为修复姿态的矫正信号。
- 在标准基准上显著提升物理合理性与全身运动估计精度,优于当前最优方法。
- 适用于需要真实感人体-环境交互的AR/VR、动作捕捉等场景。
我们提出UniCon3R,一种从单目视频进行在线人体-场景4D重建的统一前馈框架。现有前馈方法常出现身体漂浮或穿透场景的伪影,主因是缺乏对人体与环境有效交互的建模。本文目标是在推理阶段利用人体与场景间的接触信息,主动优化人体网格重建。为此,我们通过人体姿态与场景几何显式推断4D接触,并将接触作为姿态生成的修正线索。该方法可联合恢复场景几何与空间对齐的4D人体。在标准人体中心视频基准上的实验表明,UniCon3R在物理合理性与全局人体运动估计方面超越现有最优基线,同时保持快速前馈推理速度。结果验证了接触作为强大内部先验的有效性,建立了一种新的物理合理化联合重建范式。
原文摘要 · Abstract (English)
We introduce UniCon3R, a unified feed-forward framework for online human-scene 4D reconstruction from monocular video. Current feed-forward human-scene reconstruction methods suffer from artifacts, where bodies float above the ground or penetrate parts of the scene. A key reason is the lack of effective interaction modelling between the human and the environment. Our goal is to exploit contact between the human and the scene during inference to actively improve the human mesh reconstruction. To that end, we explicitly model interaction by inferring 4D contact from the human pose and scene geometry and use the contact as a corrective cue for generating the pose. This enables UniCon3R to jointly recover scene geometry and spatially aligned 4D humans within the scene. Experiments on standard human-centric video benchmarks show that UniCon3R outperforms state-of-the-art baselines on physical plausibility and global human motion estimation while preserving fast, feed-forward inference speeds. The results validate our central claim: contact serves as a powerful internal prior, thus establishing a new paradigm for physically grounded joint human-scene reconstruction. Project page is available at https://surtantheta.github.io/UniCon3R .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。