首个从动态相机视频恢复4D手部运动的方法,突破了传统单目重建的局限。
Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera

- 结合SLAM与交互手先验,实现动态视角下手部4D轨迹重建
- 在真实场景和室内数据集上显著优于现有方法,精度提升明显
- 适合需要高精度手部动作捕捉的虚拟现实与人机交互应用
我们提出Dyn-HaMR,据我们所知,首个从野外动态相机单目视频中恢复4D全局手部运动的方法。从单目视频准确重建3D手部网格对理解人类行为至关重要,广泛应用于增强现实与虚拟现实(AR/VR)。然而,现有单目手部重建方法通常依赖弱透视相机模型,仅在有限视口内模拟手部运动,难以恢复完整的3D全局轨迹,尤其在动态相机拍摄时,常导致深度估计噪声大或错误。我们的Dyn-HaMR采用多阶段、多目标优化流程,融合(i)SLAM以鲁棒估计相对相机运动,(ii)交互手先验用于生成填充并优化交互动态,确保在自遮挡下的合理性,(iii)基于先进手部追踪方法的分层初始化。在野外及室内数据集上的大量评估表明,该方法在4D全局网格恢复上显著超越现有最优方法,建立了单目动态相机视频手部运动重建的新基准。项目主页:https://dyn-hamr.github.io/
原文摘要 · Abstract (English)
We propose Dyn-HaMR, to the best of our knowledge, the first approach to reconstruct 4D global hand motion from monocular videos recorded by dynamic cameras in the wild. Reconstructing accurate 3D hand meshes from monocular videos is a crucial task for understanding human behaviour, with significant applications in augmented and virtual reality (AR/VR). However, existing methods for monocular hand reconstruction typically rely on a weak perspective camera model, which simulates hand motion within a limited camera frustum. As a result, these approaches struggle to recover the full 3D global trajectory and often produce noisy or incorrect depth estimations, particularly when the video is captured by dynamic or moving cameras, which is common in egocentric scenarios. Our Dyn-HaMR consists of a multi-stage, multi-objective optimization pipeline, that factors in (i) simultaneous localization and mapping (SLAM) to robustly estimate relative camera motion, (ii) an interacting-hand prior for generative infilling and to refine the interaction dynamics, ensuring plausible recovery under (self-)occlusions, and (iii) hierarchical initialization through a combination of state-of-the-art hand tracking methods. Through extensive evaluations on both in-the-wild and indoor datasets, we show that our approach significantly outperforms state-of-the-art methods in terms of 4D global mesh recovery. This establishes a new benchmark for hand motion reconstruction from monocular video with moving cameras. Our project page is at https://dyn-hamr.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。