从单目视频重建手部与身体动作,提升手语生成精度
DexAvatar: 3D Sign Language Reconstruction with Hand and Body Pose Priors
- 利用学习到的3D手与身体先验知识,从单目视频中重建精细动作
- 在SGNify数据集上,动作估计准确率比现有方法提升35.11%
- 适合手语生成、虚拟人动作重建等需要高精度动作还原的场景
当前手语生成多依赖大量精确的2D和3D人体姿态数据,但现有手语数据集多为视频形式,仅提供自动重建的2D关键点,缺乏准确的3D信息。且现有从手语视频中自动估计3D人体姿态的方法易受自遮挡、噪声和运动模糊影响,导致重建质量差。为此,我们提出DexAvatar框架,通过学习的3D手部与身体先验,从真实环境中的单目手语视频中重建出生物力学准确的细粒度手部关节运动与身体动作。在唯一的基准数据集SGNify上,DexAvatar相比现有最先进方法,在身体与手部姿态估计上提升35.11%。项目官网:https://github.com/kaustesseract/DexAvatar。
原文摘要 · Abstract (English)
The trend in sign language generation is centered around data-driven generative methods that require vast amounts of precise 2D and 3D human pose data to achieve an acceptable generation quality. However, currently, most sign language datasets are video-based and limited to automatically reconstructed 2D human poses (i.e., keypoints) and lack accurate 3D information. Furthermore, existing state-of-the-art for automatic 3D human pose estimation from sign language videos is prone to self-occlusion, noise, and motion blur effects, resulting in poor reconstruction quality. In response to this, we introduce DexAvatar, a novel framework to reconstruct bio-mechanically accurate fine-grained hand articulations and body movements from in-the-wild monocular sign language videos, guided by learned 3D hand and body priors. DexAvatar achieves strong performance in the SGNify motion capture dataset, the only benchmark available for this task, reaching an improvement of 35.11% in the estimation of body and hand poses compared to the state-of-the-art. The official website of this work is: https://github.com/kaustesseract/DexAvatar.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。