联合优化相机、人体姿态与3D高斯,实现野外单目真人重建
JOintGS: Joint Optimization of Cameras, Bodies and 3D Gaussians for In-the-Wild Monocular Reconstruction
- 三者协同优化,通过前景背景解耦提升精度
- 在NeuMan上提升2.1dB PSNR,实时渲染且抗初始化噪声
- 适合需要真实场景高保真人体重建的研究者
从单目RGB视频中重建高保真可驱动3D人体模型仍具挑战,尤其在无约束野外场景中,基于现成方法(如COLMAP、HMR2.0)获取的相机参数与人体姿态常不准确。尽管点云渲染(3DGS)在渲染质量与实时性上表现优异,但其高度依赖精确的相机标定与姿态标注,限制了实际应用。本文提出JOintGS,一个统一框架,通过协同优化相机外参、人体姿态与3D高斯表示,从粗初始化开始进行联合精炼。核心思想是显式分离前景与背景:静态背景高斯通过多视角一致性锚定相机估计;优化后的相机提升人体对齐精度;优化的姿态则通过静态约束消除动态伪影,改善场景重建。此外引入时序动态模块以捕捉精细姿态相关形变,并设计残差颜色场建模光照变化。在NeuMan与EMDB数据集上的大量实验表明,JOintGS在NeuMan上相较最优方法提升2.1~dB PSNR,同时保持实时渲染性能,且对噪声初始化具有显著鲁棒性。
原文摘要 · Abstract (English)
Reconstructing high-fidelity animatable 3D human avatars from monocular RGB videos remains challenging, particularly in unconstrained in-the-wild scenarios where camera parameters and human poses from off-the-shelf methods (e.g., COLMAP, HMR2.0) are often inaccurate. Splatting (3DGS) advances demonstrate impressive rendering quality and real-time performance, they critically depend on precise camera calibration and pose annotations, limiting their applicability in real-world settings. We present JOintGS, a unified framework that jointly optimizes camera extrinsics, human poses, and 3D Gaussian representations from coarse initialization through a synergistic refinement mechanism. Our key insight is that explicit foreground-background disentanglement enables mutual reinforcement: static background Gaussians anchor camera estimation via multi-view consistency; refined cameras improve human body alignment through accurate temporal correspondence; optimized human poses enhance scene reconstruction by removing dynamic artifacts from static constraints. We further introduce a temporal dynamics module to capture fine-grained pose-dependent deformations and a residual color field to model illumination variations. Extensive experiments on NeuMan and EMDB datasets demonstrate that JOintGS achieves superior reconstruction quality, with 2.1~dB PSNR improvement over state-of-the-art methods on NeuMan dataset, while maintaining real-time rendering. Notably, our method shows significantly enhanced robustness to noisy initialization compared to the baseline.Our source code is available at https://github.com/MiliLab/JOintGS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。