仅用一两视角即可生成任意姿势的动态人体3D模型
HumMorph: Generalized Dynamic Human Neural Fields from Few Views
- 从少数视角构建人体通用神经场,支持自由视角渲染
- 单视角下表现媲美顶尖方法,双视角时视觉质量显著提升
- 对噪声姿态参数更鲁棒,适合真实场景应用
我们提出HumMorph,一种新的动态人体自由视角渲染方法,支持显式姿态控制。该方法仅需少量观测视角(最少一个)即可在任意指定姿态下重建人体。模型通过前向传播实现快速推理。首先构建人体在标准T姿态下的粗略表征,融合各视角的视觉特征并利用学习到的先验知识补全缺失信息;再结合直接从观测视角提取的细粒度像素对齐特征,提供高分辨率外观信息。实验表明,在仅有一个输入视角时,HumMorph性能与当前最优方法相当;而使用两个单目视角时,视觉质量明显更优。以往泛化方法依赖同步多相机获取精确身体形状和姿态参数,而本方法在仅从观测视角估计噪声参数的更实际场景下仍表现优异,对参数误差更具鲁棒性,显著优于现有方法。
原文摘要 · Abstract (English)
We introduce HumMorph, a novel generalized approach to free-viewpoint rendering of dynamic human bodies with explicit pose control. HumMorph renders a human actor in any specified pose given a few observed views (starting from just one) in arbitrary poses. Our method enables fast inference as it relies only on feed-forward passes through the model. We first construct a coarse representation of the actor in the canonical T-pose, which combines visual features from individual partial observations and fills missing information using learned prior knowledge. The coarse representation is complemented by fine-grained pixel-aligned features extracted directly from the observed views, which provide high-resolution appearance information. We show that HumMorph is competitive with the state-of-the-art when only a single input view is available, however, we achieve results with significantly better visual quality given just 2 monocular observations. Moreover, previous generalized methods assume access to accurate body shape and pose parameters obtained using synchronized multi-camera setups. In contrast, we consider a more practical scenario where these body parameters are noisily estimated directly from the observed views. Our experimental results demonstrate that our architecture is more robust to errors in the noisy parameters and clearly outperforms the state of the art in this setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。