用野外手部图像预训练3D手姿估计,提升精度
Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
- 基于野外视频提取200万张手部图像,用对比学习对齐相似手姿
- 在FreiHand等数据集上提升15%、10%、4%,超越现有方法
- 适合做手部姿态估计的开发者,尤其关注数据多样性
我们提出一种基于野外手部图像的对比学习框架HandCLR,用于3D手部姿态估计器的预训练。尽管大规模图像预训练已取得良好效果,但以往的3D手部姿态预训练方法未充分利用来自野外视频的丰富手部图像。为实现可扩展的预训练,我们从近期以人为中心的视频(如100DOH和Ego4D)中收集了超过200万张手部图像,并设计了一种新的对比学习方法。该方法聚焦于手部姿态的相似性,将不同样本中相似的手姿对在隐空间中拉近。实验表明,该方法优于仅通过单图增强生成正样本的传统对比学习方法。在多个数据集上显著超越当前最优方法:FreiHand提升15%,DexYCB提升10%,AssemblyHands提升4%。
原文摘要 · Abstract (English)
We present a contrastive learning framework based on in-the-wild hand images tailored for pre-training 3D hand pose estimators, dubbed HandCLR. Pre-training on large-scale images achieves promising results in various tasks, but prior 3D hand pose pre-training methods have not fully utilized the potential of diverse hand images accessible from in-the-wild videos. To facilitate scalable pre-training, we first prepare an extensive pool of hand images from in-the-wild videos and design our method with contrastive learning. Specifically, we collected over 2.0M hand images from recent human-centric videos, such as 100DOH and Ego4D. To extract discriminative information from these images, we focus on the similarity of hands; pairs of similar hand poses originating from different samples, and propose a novel contrastive learning method that embeds similar hand pairs closer in the latent space. Our experiments demonstrate that our method outperforms conventional contrastive learning approaches that produce positive pairs sorely from a single image with data augmentation. We achieve significant improvements over the state-of-the-art method in various datasets, with gains of 15% on FreiHand, 10% on DexYCB, and 4% on AssemblyHands.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。