用多尺度贴图与扩散修复重建被遮挡的人体,效果更完整连贯。
InpaintHuman: Reconstructing Occluded Humans with Multi-Scale UV Mapping and Identity-Preserving Diffusion Inpainting
- 设计多尺度UV贴图,分层插值恢复遮挡区域几何细节。
- 引入身份保持的扩散修复模块,实现时序一致的高保真重建。
- 适合需要高质量人体重建与动画的应用场景。
从单目视频中重建完整且可动画化的3D人体形象仍具挑战性,尤其在严重遮挡情况下。尽管3D高斯点云已实现逼真人体渲染,现有方法在观测不全时仍易产生几何畸变和时间不一致问题。本文提出InpaintHuman,一种从遮挡单目视频生成高保真、完整且可动画化人体形象的新方法。核心创新包括:(i) 多尺度UV参数化表示结合层次化粗到细特征插值,有效恢复遮挡区域并保留几何细节;(ii) 身份保持的扩散修复模块,融合文本反演与语义条件引导,实现个体特定、时序一致的完成。不同于基于SDS的方法,本方案采用像素级直接监督以保障身份一致性。在合成基准(PeopleSnapshot、ZJU-MoCap)和真实场景(OcMotion)上的实验表明,该方法在多种姿态与视角下均取得有竞争力的表现,重建质量持续提升。
原文摘要 · Abstract (English)
Reconstructing complete and animatable 3D human avatars from monocular videos remains challenging, particularly under severe occlusions. While 3D Gaussian Splatting has enabled photorealistic human rendering, existing methods struggle with incomplete observations, often producing corrupted geometry and temporal inconsistencies. We present InpaintHuman, a novel method for generating high-fidelity, complete, and animatable avatars from occluded monocular videos. Our approach introduces two key innovations: (i) a multi-scale UV-parameterized representation with hierarchical coarse-to-fine feature interpolation, enabling robust reconstruction of occluded regions while preserving geometric details; and (ii) an identity-preserving diffusion inpainting module that integrates textual inversion with semantic-conditioned guidance for subject-specific, temporally coherent completion. Unlike SDS-based methods, our approach employs direct pixel-level supervision to ensure identity fidelity. Experiments on synthetic benchmarks (PeopleSnapshot, ZJU-MoCap) and real-world scenarios (OcMotion) demonstrate competitive performance with consistent improvements in reconstruction quality across diverse poses and viewpoints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。