用稀疏低重叠摄像头重建4D人体场景,效果优于现有方法
4D Human-Scene Reconstruction from Low-Overlap Captures

- 分离背景与人体,分步优化提升重建精度
- 通过视频扩散模型生成数百个新视角,增强背景监督
- 适合需要高质量动态人体重建的影视与虚拟制作场景
现有动态人体体积捕获依赖密集相机阵列以实现高保真度。但在真实场景中,仅有少量低重叠摄像头可用,导致输出质量下降且存在大片未观测区域。近期4D重建方法虽聚焦于低重叠设置,但仍会在欠观测区域产生明显伪影。视频扩散模型虽为替代方案,但对人体几何重建存在不一致问题。为此,我们提出StudioRecon,一种从稀疏、低重叠摄像头重建4D人体场景的流程。通过解耦背景与人体,利用视频扩散模型合成数百个受控新视角以增强背景监督;通过跨视角身份关联与多视图关键点三角化,鲁棒初始化可变形高斯人体;最后,采用运动自适应一致性注入的递归增强模块,进一步消除残余伪影。我们在四个真实数据集上实现最佳新视角合成效果,并展示了新轨迹渲染与人体替换等应用。
原文摘要 · Abstract (English)
Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in real-world scenarios, only a handful of low-overlap cameras are available, which degrades the output quality and leaves large areas unobserved. Recent 4D reconstruction methods have focused on low-overlap settings, yet they still produce noticeable artifacts in under-observed regions. Video diffusion models have emerged as another option, but they show geometrically inconsistent results for humans. To address these limitations, we propose StudioRecon, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans. We densify background supervision by synthesizing hundreds of camera-controlled novel views with a video diffusion model. We also robustly initialize deformable Gaussian humans with cross-view identity association and triangulated multi-view keypoint fitting. Finally, our recursive enhancement module with motion-adaptive consistency injection harmonizes the composed output, thereby further avoiding remaining artifacts. We achieve state-of-the-art novel view synthesis across four real-world datasets and demonstrate applications such as novel trajectory rendering and human replacement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。