用社交互动信息提升第一人称视角人体三维姿态估计精度。
Social EgoMesh Estimation
- 仅用潜在扩散模型结合场景与人际互动估计佩戴者3D姿态。
- 相比现有最佳方法,姿态误差(MPJPE)降低53%。
- 首次揭示距离和视线方向对估计精度的影响,适合虚拟现实研究者。
在第一人称视频中准确估计佩戴者三维姿态对虚拟与增强现实中的行为建模至关重要。由于头戴前向摄像头导致身体可见性受限,该任务面临独特挑战。现有研究利用场景和自我运动信息,却忽略了人类的交互特性。本文提出社交第一人称姿态网格估计框架(SEE-ME),首次仅通过潜变量概率扩散模型,结合场景及佩戴者与交互对象的社会互动信息进行姿态估计。深入研究表明,人际距离与视线方向对估计精度影响显著。总体而言,SEE-ME将姿态估计误差(MPJPE)降低了53%,优于当前最优方法。代码已公开于https://github.com/L-Scofano/SEEME。
原文摘要 · Abstract (English)
Accurately estimating the 3D pose of the camera wearer in egocentric video sequences is crucial to modeling human behavior in virtual and augmented reality applications. The task presents unique challenges due to the limited visibility of the user's body caused by the front-facing camera mounted on their head. Recent research has explored the utilization of the scene and ego-motion, but it has overlooked humans' interactive nature. We propose a novel framework for Social Egocentric Estimation of body MEshes (SEE-ME). Our approach is the first to estimate the wearer's mesh using only a latent probabilistic diffusion model, which we condition on the scene and, for the first time, on the social wearer-interactee interactions. Our in-depth study sheds light on when social interaction matters most for ego-mesh estimation; it quantifies the impact of interpersonal distance and gaze direction. Overall, SEE-ME surpasses the current best technique, reducing the pose estimation error (MPJPE) by 53%. The code is available at https://github.com/L-Scofano/SEEME.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。