单目图像估3D手姿,无需相机参数也能准
Monocular 3D Hand Pose Estimation with Implicit Camera Alignment
- 用关键点对齐+指尖损失隐式校准相机参数
- 在EgoDexter和Dexter+Object上达顶尖性能
- 适合无相机信息的野外场景应用
从单张彩色图像估计3D手部姿态是增强现实(AR)、虚拟现实(VR)、人机交互(HCI)和机器人领域的重要问题。除了缺乏深度信息外,遮挡、关节结构复杂以及需要已知相机参数也带来了挑战。本文提出一种优化流程,基于2D关键点输入估计3D手部姿态,包含关键点对齐步骤和指尖损失项,以避免显式知晓或估计相机参数。我们在EgoDexter和Dexter+Object基准上评估该方法,结果表明其性能可与当前最优方法媲美,并在处理无先验相机知识的“真实场景”图像时表现出强鲁棒性。定量分析显示,尽管使用了手部先验,2D关键点估计精度仍对最终结果有显著影响。代码已公开于项目页面 https://cpantazop.github.io/HandRepo/
原文摘要 · Abstract (English)
Estimating the 3D hand articulation from a single color image is an important problem with applications in Augmented Reality (AR), Virtual Reality (VR), Human-Computer Interaction (HCI), and robotics. Apart from the absence of depth information, occlusions, articulation complexity, and the need for camera parameters knowledge pose additional challenges. In this work, we propose an optimization pipeline for estimating the 3D hand articulation from 2D keypoint input, which includes a keypoint alignment step and a fingertip loss to overcome the need to know or estimate the camera parameters. We evaluate our approach on the EgoDexter and Dexter+Object benchmarks to showcase that it performs competitively with the state-of-the-art, while also demonstrating its robustness when processing "in-the-wild" images without any prior camera knowledge. Our quantitative analysis highlights the sensitivity of the 2D keypoint estimation accuracy, despite the use of hand priors. Code is available at the project page https://cpantazop.github.io/HandRepo/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。