arXiv:2410.03665cs.CVcs.AI2024-10CVPR被引 56

仅用头戴设备数据,实时估算人体与手部三维动作。

Estimating Body and Hand Motion in an Ego-sensed World

论文配图:Estimating Body and Hand Motion in an Ego-sensed World
图 1 · 摘自论文原文
  • 基于条件扩散模型,利用视角内运动信息生成姿态
  • 头部运动参数化使估计误差降低18%,手部估计精度提升40%
  • 适合增强现实、可穿戴计算等场景的实时动作捕捉

我们提出EgoAllo系统,仅通过头戴设备的自视角SLAM位姿和图像,实现对佩戴者3D身体姿态、身高及手部参数的估计,结果在场景的共坐标系中呈现。核心思路在于表征设计:提出空间与时间不变性准则,推导出头部运动条件化参数化方法,使估计性能最高提升18%。同时,系统通过人体运动学与时间约束,显著改善手部估计,使单帧世界坐标误差降低40%。项目主页:https://egoallo.github.io/

原文摘要 · Abstract (English)

We present EgoAllo, a system for human motion estimation from a head-mounted device. Using only egocentric SLAM poses and images, EgoAllo guides sampling from a conditional diffusion model to estimate 3D body pose, height, and hand parameters that capture a device wearer's actions in the allocentric coordinate frame of the scene. To achieve this, our key insight is in representation: we propose spatial and temporal invariance criteria for improving model performance, from which we derive a head motion conditioning parameterization that improves estimation by up to 18%. We also show how the bodies estimated by our system can improve hand estimation: the resulting kinematic and temporal constraints can reduce world-frame errors in single-frame estimates by 40%. Project page: https://egoallo.github.io/

动作估计头戴设备扩散模型人体姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。