多视角第一人称手势追踪新方法,提升虚拟现实交互精度
1st Place Solution of Multiview Egocentric Hand Tracking Challenge ECCV2024
- 融合多视角图像与相机参数,联合估计手形与姿态
- 在Umetrack上达13.92mm MPJPE,HOT3D上达21.66mm MPJPE
- 采用裁剪抖动和神经平滑后处理,适合高精度手势应用
多视角第一人称手势追踪是一项挑战性任务,在虚拟现实交互中至关重要。本文提出一种方法,利用多视角输入图像和相机外参,联合估计手形与姿态。为降低对相机布局的过拟合,引入裁剪抖动和外参噪声增强。此外,提出一种离线神经平滑后处理方法,进一步提升手部位置与姿态的准确性。该方法在Umetrack数据集上达到13.92mm MPJPE,HOT3D数据集上达到21.66mm MPJPE。
原文摘要 · Abstract (English)
Multi-view egocentric hand tracking is a challenging task and plays a critical role in VR interaction. In this report, we present a method that uses multi-view input images and camera extrinsic parameters to estimate both hand shape and pose. To reduce overfitting to the camera layout, we apply crop jittering and extrinsic parameter noise augmentation. Additionally, we propose an offline neural smoothing post-processing method to further improve the accuracy of hand position and pose. Our method achieves 13.92mm MPJPE on the Umetrack dataset and 21.66mm MPJPE on the HOT3D dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。