arXiv:2508.08798cs.CV2025-08被引 1

单目视频下人体重建,用分块神经辐射场实现更自然的形变与遮挡恢复。

MonoPartNeRF:Human Reconstruction from Monocular Video via Part-Based Neural Radiance Fields

  • 分块姿态嵌入+双向形变模型,实现身体各部位灵活动态建模。
  • 在ZJU-MoCap和MonoCap数据集上,关节对齐与纹理保真度显著提升。
  • 适合关注单视角动态人体渲染、形变连续性的研究者。

近年来,神经辐射场(NeRF)在动态人体重建与渲染方面取得显著进展。基于人体分割的分块渲染范式可根据结构复杂度灵活分配参数,提升表示效率。然而,现有方法在复杂姿态变化下仍存在部分边界过渡不自然、单目条件下遮挡区域重建不准等问题。本文提出MonoPartNeRF,一种新颖的单目动态人体渲染框架,确保形变过渡平滑并实现鲁棒的遮挡恢复。首先,构建结合刚性与非刚性变换的双向形变模型,建立观测空间与标准空间之间的连续可逆映射;采样点被投影至参数化表面-时间空间(u, v, t),更好捕捉非刚性运动;一致性损失进一步抑制形变带来的伪影与不连续。引入基于部位的姿势嵌入机制,将全局姿态向量分解为局部关节嵌入,结合关键帧姿态检索与三轴插值,引导姿态感知特征采样。通过注意力机制集成可学习外观码,有效建模动态纹理变化。在ZJU-MoCap和MonoCap数据集上的实验表明,本方法在复杂姿态与遮挡条件下显著优于现有方法,实现了更优的关节对齐、纹理保真度与结构连续性。

原文摘要 · Abstract (English)

In recent years, Neural Radiance Fields (NeRF) have achieved remarkable progress in dynamic human reconstruction and rendering. Part-based rendering paradigms, guided by human segmentation, allow for flexible parameter allocation based on structural complexity, thereby enhancing representational efficiency. However, existing methods still struggle with complex pose variations, often producing unnatural transitions at part boundaries and failing to reconstruct occluded regions accurately in monocular settings. We propose MonoPartNeRF, a novel framework for monocular dynamic human rendering that ensures smooth transitions and robust occlusion recovery. First, we build a bidirectional deformation model that combines rigid and non-rigid transformations to establish a continuous, reversible mapping between observation and canonical spaces. Sampling points are projected into a parameterized surface-time space (u, v, t) to better capture non-rigid motion. A consistency loss further suppresses deformation-induced artifacts and discontinuities. We introduce a part-based pose embedding mechanism that decomposes global pose vectors into local joint embeddings based on body regions. This is combined with keyframe pose retrieval and interpolation, along three orthogonal directions, to guide pose-aware feature sampling. A learnable appearance code is integrated via attention to model dynamic texture changes effectively. Experiments on the ZJU-MoCap and MonoCap datasets demonstrate that our method significantly outperforms prior approaches under complex pose and occlusion conditions, achieving superior joint alignment, texture fidelity, and structural continuity.

人体重建神经辐射场单目视频分块建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。