arXiv:2606.19156cs.CV2026-06

首个端到端重建第一人称视频中动态手部的4D模型,速度快且泛化强。

Hand-4DGS: Feed-Forward 3D Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos

论文配图:Hand-4DGS: Feed-Forward 3D Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos
图 1 · 摘自论文原文
  • 用网格引导表示和时序卷积建模手部结构与动态变化
  • 在H2O和ARCTIC数据集上显著优于基线方法,推理速度达60 FPS
  • 无需3D手部姿态标注,依赖2D图像监督即可实现良好泛化

从第一人称视频中进行动态3D手部重建对下一代计算平台(如AR/VR、AI眼镜)至关重要。尽管意义重大,现有工作多集中于多视角3D手部重建或4D人体重建。由于头部快速运动、手部动态剧烈、严重遮挡以及单视角观测的固有模糊性,第一人称4D手部重建仍具挑战。为此,我们提出Hand-4DGS,首个直接从第一人称视频中重建动态4D手部的端到端框架,支持高速(~60 FPS)推理和强泛化能力。方法结合网格引导表示引入结构先验,并采用时序卷积建模动态运动。我们在两个具有挑战性的第一人称数据集H2O和ARCTIC上评估,结果显著优于基线。该方法得益于端到端网络的泛化能力,以及通过高斯点阵实现的有效2D图像监督,无需昂贵的3D手部姿态真值标注。

原文摘要 · Abstract (English)

Dynamic 3D hand reconstruction from egocentric videos is essential for next-generation computing platforms such as AR/VR and AI glasses. Despite its importance, most prior works focus either on multi-view 3D hand reconstruction or on 4D human body reconstruction. Egocentric 4D hand reconstruction remains challenging due to fast head motion, rapid hand dynamics, severe occlusions, and inherent ambiguity from single-view observations. To address these challenges, we introduce Hand-4DGS, the first feed-forward framework for reconstructing dynamic 4D hands directly from egocentric videos, enabling both fast (~60 FPS) inference and strong generalization. Our approach incorporates a mesh-guided representation for structural priors and temporal convolutions to model dynamic motion. We evaluate our framework on two challenging egocentric datasets, H2O and ARCTIC, and demonstrate significant improvements over baselines. Our method benefits from the generalization capability of feed-forward networks and effective 2D image supervision through Gaussian splatting, without requiring expensive 3D hand pose ground-truth annotations.

4D重建手部建模端到端第一人称视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。