用时空变压器提升单目穿衣人体三维重建的细节与稳定性
Disambiguating Monocular Reconstruction of 3D Clothed Human with Spatial-Temporal Transformer
- 引入空间-时间变压器,分别捕捉全局纹理和帧间时序特征
- 在Adobe和MonoPerfCap数据集上优于现有方法,低光环境下仍稳定
- 适合关注人体三维重建、动作捕捉与视觉算法优化的研究者
从单目摄像头数据中重建三维穿衣人体极具挑战,主要源于视角限制和图像模糊。尽管基于隐式函数的方法结合参数化模型已取得进展,但仍存在两个问题:一是因视角不可见导致人体背面细节模糊,依赖卷积神经网络预测的背面法线图过于平滑;二是单帧图像受光照和身体运动影响存在局部模糊,而隐式函数对像素变化敏感。为此,本文提出空间-时间变压器(STT)网络。空间变压器用于提取全局信息以生成更准确的法线图,建立整体纹理与形状关联;时间变压器则利用相邻帧提取时序特征,增强隐式网络输入的准确性。同时,通过联合标记建立帧间局部对应关系,进一步提升时序特征精度。在Adobe和MonoPerfCap数据集上的实验表明,本方法优于当前最优方法,并在低光户外条件下保持鲁棒性。
原文摘要 · Abstract (English)
Reconstructing 3D clothed humans from monocular camera data is highly challenging due to viewpoint limitations and image ambiguity. While implicit function-based approaches, combined with prior knowledge from parametric models, have made significant progress, there are still two notable problems. Firstly, the back details of human models are ambiguous due to viewpoint invisibility. The quality of the back details depends on the back normal map predicted by a convolutional neural network (CNN). However, the CNN lacks global information awareness for comprehending the back texture, resulting in excessively smooth back details. Secondly, a single image suffers from local ambiguity due to lighting conditions and body movement. However, implicit functions are highly sensitive to pixel variations in ambiguous regions. To address these ambiguities, we propose the Spatial-Temporal Transformer (STT) network for 3D clothed human reconstruction. A spatial transformer is employed to extract global information for normal map prediction. The establishment of global correlations facilitates the network in comprehending the holistic texture and shape of the human body. Simultaneously, to compensate for local ambiguity in images, a temporal transformer is utilized to extract temporal features from adjacent frames. The incorporation of temporal features can enhance the accuracy of input features in implicit networks. Furthermore, to obtain more accurate temporal features, joint tokens are employed to establish local correspondences between frames. Experimental results on the Adobe and MonoPerfCap datasets have shown that our method outperforms state-of-the-art methods and maintains robust generalization even under low-light outdoor conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。