融合视觉与手套陀螺仪数据,提升遮挡下手部三维追踪精度。
AVI-HT: Adaptive Vision-IMU Fusion for 3D Hand Tracking

- 通过跨传感器注意力机制动态调节视觉与陀螺仪信任度。
- 在10万+样本数据集上,关键点误差降低16.1%~24.2%。
- 适合手势交互、遮挡严重等复杂场景的高精度追踪应用。
我们提出AVI-HT,一种自适应视觉-惯性(IMU)融合方法,用于通过联合建模佩戴式6-自由度惯性传感器信号与第一人称视觉图像,实现3D手部姿态追踪。该方法在手-物体交互(HOI)场景中,尤其在严重视觉遮挡情况下,显著提升了追踪精度与可用性。其成功依赖于两个互补因素:(1) 基于动作捕捉系统生成的真实3D手部姿态标签,对身体上同步的多模态视觉-惯性数据流进行配对训练;(2) 一种跨传感器深度注意力机制,可自适应地调节对视觉和各惯性传感器的信任权重。为评估真实场景性能,我们在包含10万+对齐样本的DexGloveHOI数据集上进行了广泛实验,涵盖日常任务中的多种物体操作。在两种手部模型(UmeTrack, MANO)下,与多种单模态与多模态追踪方法对比,结果显示AVI-HT将平均关键点误差降低16.1%,其腕部对齐变体进一步降低24.2%。消融实验揭示了不同活动类型下各手指惯性传感器的贡献差异,并分析了模型对惯性噪声和视觉-惯性时间错位的敏感性。
原文摘要 · Abstract (English)
We present AVI-HT, an adaptive visual-IMU fusion approach for tracking 3D hand poses by jointly modeling the egocentric image with on-glove 6-DoF IMU signals. AVI-HT achieves significantly improved accuracy and availability, particularly in hand-object interaction (HOI) scenarios involving heavy visual occlusion. Two complementary ingredients underpin its success: (1) synchronized multi-modal training data pairing on-body vision-IMU sensor streams with ground-truth 3D hand poses from a motion-capture system, and (2) a cross-sensor deep attention mechanism that adaptively modulates the trust assigned to the vision and individual IMU sensors. To evaluate AVI-HT in real-world settings, we conduct extensive experiments on our DexGloveHOI dataset that consists of 100K+ pairwise vision-IMU samples with synchronized 3D annotated poses, in which users manipulate a variety of objects during daily tasks. We compare against multiple single- and multi-modal tracking approaches under two hand models (UmeTrack, MANO). The results show that AVI-HT reduces mean keypoint error by 16.1% and its wrist-aligned variant by 24.2% over the baselines. Ablation studies further reveal the per-finger contribution of IMU sensors across activity types, and the model's sensitivity to IMU noise and temporal misalignment in vision-IMU fusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。