arXiv:2605.12498cs.CVcs.GR2026-05International Conf…被引 2

单目相机下精准还原手部3D姿态,不依赖特定设备

EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

论文配图:EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera
图 1 · 摘自论文原文
  • 用可微分前臂建模稳定手部姿态,统一处理多种镜头畸变
  • 在HOT3D上相机空间MPJPE降低28%,跨设备性能一致
  • 适合AR/VR、远程交互等需要轻量无感传感的场景

从单个头戴式相机的用户视角重建绝对3D手部姿态与形状,对AR/VR、远程呈现及以手为中心的操作任务至关重要,要求感知系统紧凑且无感。尽管单目RGB方法已有进展,但仍受深度尺度模糊限制,难以泛化至不同头戴设备的光学配置。因此,模型通常需在特定设备数据集上大量训练,成本高且耗时。本文提出EgoForce,一种单目3D手部重建框架,能从用户视角(相机空间)恢复鲁棒的绝对3D手部姿态与位置。该方法通过单一统一网络支持鱼眼、透视及广角畸变相机模型。其核心包括:可微分的前臂表示以稳定手部姿态;统一的臂-手变换器,从单个视图预测手部与前臂几何结构,缓解深度尺度模糊;以及射线空间闭式求解器,实现跨多种头戴相机模型的绝对3D姿态恢复。在三个头戴式基准上的实验表明,EgoForce达到当前最优3D精度,在HOT3D数据集上相机空间MPJPE降低28%以上,并保持跨相机配置的一致性能。

原文摘要 · Abstract (English)

Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulation tasks, where sensing must remain compact and unobtrusive. While monocular RGB methods have made progress, they remain constrained by depth-scale ambiguity and struggle to generalize across the diverse optical configurations of head-mounted devices. As a result, models typically require extensive training on device-specific datasets, which are costly and laborious to acquire. This paper addresses these challenges by introducing EgoForce, a monocular 3D hand reconstruction framework that recovers robust, absolute 3D hand pose and its position from the user's (camera-space) viewpoint. EgoForce operates across fisheye, perspective, and distorted wide-FOV camera models using a single unified network. Our approach combines a differentiable forearm representation that stabilizes hand pose, a unified arm-hand transformer that predicts both hand and forearm geometry from a single egocentric view, mitigating depth-scale ambiguity, and a ray space closed-form solver that enables absolute 3D pose recovery across diverse head-mounted camera models. Experiments on three egocentric benchmarks show that EgoForce achieves state-of-the-art 3D accuracy, reducing camera-space MPJPE by up to 28% on the HOT3D dataset compared to prior methods and maintaining consistent performance across camera configurations. For more details, visit the project page at https://dfki-av.github.io/EgoForce.

3D手姿单目重建AR/VR相机空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。