端到端实现野外复杂场景下多手3D定位与重建,效率精度双提升。
WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
- 基于全卷积网络实时定位手部,结合Transformer高保真重建3D姿态。
- 在200万张野外手部图像上训练,覆盖多种光照遮挡条件。
- 无需时间信息即可实现单目视频流畅3D手部追踪,适合真实场景应用。
近年来,3D手部姿态估计因在人机交互、虚拟现实和机器人领域的广泛应用而备受关注。相比之下,手部检测流程仍存在明显短板,严重制约了真实世界多手重建系统的构建。本文提出一种数据驱动的端到端多手重建流水线,包含实时全卷积手部定位模块与高保真Transformer-based 3D手部重建模型。为克服以往方法局限并构建鲁棒稳定的检测网络,我们引入一个包含超过200万张野外手部图像的大规模数据集,涵盖多样化的光照、照明与遮挡条件。所提方法在主流2D与3D基准测试中均优于现有方法,在效率与精度上表现更优。最后,我们展示了该流水线在无时序依赖情况下,仅通过单目视频实现平滑3D手部追踪的有效性。代码、模型与数据集已公开:https://rolpotamias.github.io/WiLoR。
原文摘要 · Abstract (English)
In recent years, 3D hand pose estimation methods have garnered significant attention due to their extensive applications in human-computer interaction, virtual reality, and robotics. In contrast, there has been a notable gap in hand detection pipelines, posing significant challenges in constructing effective real-world multi-hand reconstruction systems. In this work, we present a data-driven pipeline for efficient multi-hand reconstruction in the wild. The proposed pipeline is composed of two components: a real-time fully convolutional hand localization and a high-fidelity transformer-based 3D hand reconstruction model. To tackle the limitations of previous methods and build a robust and stable detection network, we introduce a large-scale dataset with over than 2M in-the-wild hand images with diverse lighting, illumination, and occlusion conditions. Our approach outperforms previous methods in both efficiency and accuracy on popular 2D and 3D benchmarks. Finally, we showcase the effectiveness of our pipeline to achieve smooth 3D hand tracking from monocular videos, without utilizing any temporal components. Code, models, and dataset are available https://rolpotamias.github.io/WiLoR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。