arXiv:2607.29633cs.CV2026-07中稿 · ACM Multimedia 202…

用3D高斯点云实现单图手部三维重建,克服遮挡难题。

OASIS: Occlusion-aware Single-image Hand Avatar Reconstruction via 3D Gaussian Splatting

论文配图:OASIS: Occlusion-aware Single-image Hand Avatar Reconstruction via 3D Gaussian Splatting
图 1 · 摘自论文原文
  • 通过几何对齐视觉特征点,提升单视图细节还原能力。
  • 引入可见性感知注意力机制,有效处理严重自遮挡问题。
  • 适合手部动画、虚拟试穿等需要高精度手部建模的应用。

单图像3D手部虚拟人重建因自遮挡严重且手部关节变形复杂而极具挑战。现有方法多依赖隐式NeRF类表示,计算成本高且难以保留精细结构。本文提出OASIS框架,基于3D高斯点云实现高效重建。通过将输入图像观测与3D手部几何显式对齐,并上下文自适应地提取视觉特征点,增强图像线索的表达。针对遮挡导致的视觉可靠性下降问题,设计可见性条件下的点-图像注意力机制,确保关键特征可靠传递至几何点。为捕捉非刚性形变,引入基于网格的特征表示,使高斯点变形受局部表面拉伸引导。采用单次适配策略,先从多身份数据学习通用手部先验,再适配目标图像进行定制化重建。大量实验表明,OASIS在复杂姿态和真实场景下均优于现有基线,在视觉保真度与效率上表现突出,并在文本生成手部、材质编辑等下游任务中展现出强泛化能力。

原文摘要 · Abstract (English)

Single-image 3D hand avatar reconstruction is fundamentally ill-posed and particularly challenging due to limited visual evidence under severe self-occlusion and the complex pose-dependent deformation of highly articulated hands. Existing methods predominantly rely on implicit NeRF-style representations, whose volumetric fitting is computationally expensive and often struggles to preserve fine-grained hand details. In this work, we present OASIS, a tailored 3D Gaussian Splatting framework for single-image hand avatar reconstruction. To faithfully encode sparse image-specific appearance cues in single-view reconstruction, we construct geometry-aligned visual evidence tokens by explicitly aligning input image observations with 3D hand geometry and context-adaptively tokenizing the resulting visual evidence. Since severe self-occlusion makes the reliability of image evidence inherently visibility-dependent, we introduce a visibility-conditioned point-image attention to reliably transfer visual evidence to geometric tokens, yielding occlusion-aware Gaussian features for faithful and robust reconstruction. To further capture non-rigid deformation of articulated hands, we introduce a Feature-on-Mesh representation to enable Gaussian deformation to be guided by local surface stretching. Under this framework, we adopt a one-shot adaptation scheme that learns a shared hand prior from multi-identity training data and then fits it to a target image for target-specific reconstruction. Extensive experiments show that OASIS outperforms existing baselines in both visual fidelity and efficiency across challenging poses and in-the-wild scenarios, and further demonstrates strong versatility in downstream applications such as text-to-avatar generation and texture editing.

手部重建3D高斯遮挡处理虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。