arXiv:2602.05190cs.CVcs.GR2026-02

用姿态引导高保真人体新视角合成,实时渲染且更稳定。

PoseGaussian: Pose-Driven Novel View Synthesis for Robust 3D Human Reconstruction

  • 姿态同时用于优化深度和保持时序一致性
  • 在三个数据集上实现PSNR 30.86、SSIM 0.979的高质量重建
  • 适合需要实时动态人体建模的应用场景

我们提出PoseGaussian,一种基于姿态引导的高斯点云渲染框架,用于高保真人体新视角合成。人体姿态在设计中扮演双重角色:作为结构先验,与颜色编码器融合以优化深度估计;作为时序线索,通过专用姿态编码器提升帧间一致性。这些组件整合为一个全可微、端到端可训练的流程。与以往仅将姿态用作条件或用于变形的方法不同,PoseGaussian将姿态信号嵌入几何与时间两个阶段,显著提升鲁棒性与泛化能力。该方法专为解决动态人体场景中的关节运动与严重自遮挡问题而设计。值得注意的是,系统实现实时渲染,达到100 FPS,保持标准高斯点云管道的效率。我们在ZJU-MoCap、THuman2.0及自建数据集上验证方法,结果表明其在感知质量与结构准确性方面达到当前最优水平(PSNR 30.86,SSIM 0.979,LPIPS 0.028)。

原文摘要 · Abstract (English)

We propose PoseGaussian, a pose-guided Gaussian Splatting framework for high-fidelity human novel view synthesis. Human body pose serves a dual purpose in our design: as a structural prior, it is fused with a color encoder to refine depth estimation; as a temporal cue, it is processed by a dedicated pose encoder to enhance temporal consistency across frames. These components are integrated into a fully differentiable, end-to-end trainable pipeline. Unlike prior works that use pose only as a condition or for warping, PoseGaussian embeds pose signals into both geometric and temporal stages to improve robustness and generalization. It is specifically designed to address challenges inherent in dynamic human scenes, such as articulated motion and severe self-occlusion. Notably, our framework achieves real-time rendering at 100 FPS, maintaining the efficiency of standard Gaussian Splatting pipelines. We validate our approach on ZJU-MoCap, THuman2.0, and in-house datasets, demonstrating state-of-the-art performance in perceptual quality and structural accuracy (PSNR 30.86, SSIM 0.979, LPIPS 0.028).

人体重建新视角合成高斯溅射姿态引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。