arXiv:2508.09629cs.CV2025-08中稿 · WACV 2026

用纹理信息提升单目手部三维重建的精度与真实感

Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors

  • 将像素级纹理映射到UV空间,实现图像与三维手形的密集对齐
  • 在HaMeR模型上加入纹理模块后,姿态与形状估计更准确
  • 轻量模块可无缝接入现有重建流程,适合追求真实感的视觉任务

我们重新审视纹理在单目三维手部重建中的作用,不再仅将其视为提升真实感的附加项,而是作为密集且空间定位明确的监督信号,主动支持姿态与形状估计。观察发现,即使高性能模型中,预测的手部几何与图像外观之间的重叠仍不理想,表明纹理对齐可能被低估。为此,我们提出一个轻量级纹理模块,将逐像素观测嵌入UV纹理空间,并实现预测与观测手部外观间的新型密集对齐损失。该方法依赖可微渲染管线及已知拓扑的图像到3D手部网格映射模型,支持将带纹理的手部反投影至图像并进行像素级对齐。模块自包含且易于集成至现有重建流程。为凸显纹理监督的价值,我们在高精度但无纹理增强的Transformer架构HaMeR基础上进行扩展,结果在准确性和真实性上均有提升,验证了外观引导对齐在手部重建中的有效性。

原文摘要 · Abstract (English)

We revisit the role of texture in monocular 3D hand reconstruction, not as an afterthought for photorealism, but as a dense, spatially grounded cue that can actively support pose and shape estimation. Our observation is simple: even in high-performing models, the overlay between predicted hand geometry and image appearance is often imperfect, suggesting that texture alignment may be an underused supervisory signal. We propose a lightweight texture module that embeds per-pixel observations into UV texture space and enables a novel dense alignment loss between predicted and observed hand appearances. Our approach assumes access to a differentiable rendering pipeline and a model that maps images to 3D hand meshes with known topology, allowing us to back-project a textured hand onto the image and perform pixel-based alignment. The module is self-contained and easily pluggable into existing reconstruction pipelines. To isolate and highlight the value of texture-guided supervision, we augment HaMeR, a high-performing yet unadorned transformer architecture for 3D hand pose estimation. The resulting system improves both accuracy and realism, demonstrating the value of appearance-guided alignment in hand reconstruction.

三维重建纹理对齐手部建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。