用扩散模型融合第一人称视角与第三人称观察,实现多人交互时的精准3D姿态估计。
Everybody Tracking Every Body

- 基于头动推断姿态,结合高可靠性的第三人称观测进行融合
- 在多人体数据集上相比纯运动或纯视觉基线提升精度
- 适用于智能眼镜等可穿戴设备中的多人行为感知场景
我们解决从具有中心化协调的自我中心视角中对多名互动人员进行3D身体姿态估计的问题。每个人佩戴摄像头和惯性测量单元(IMU),通过视觉惯性里程计(VIO SLAM)对每个自我中心摄像头在空间中的运动进行高精度跟踪。一个人的第一人称视角能提供其他人的第三人称观测,但这些外视角观测稀疏、间断且可靠性高度可变,因摄像头与目标主体均在移动。为整合这些同步的数据流,我们提出一种基于扩散模型的方法,将基于自我中心摄像头运动推断出的头部运动姿态估计与第三人称姿态观测相融合,并同时考虑观测内容与可靠性。该模型在单人动作捕捉数据与多人视频混合数据上进行训练,以学习丰富的身体运动轨迹先验及视频观测可靠性先验。在具有挑战性的多人数据集上的评估表明,我们的融合方法在绝对和相对姿态精度上均优于仅依赖运动或仅依赖视觉的基线方法。
原文摘要 · Abstract (English)
We address the problem of 3D body pose estimation of multiple interacting people from their egocentric views with centralized coordination. Each individual wears a camera recording egocentric video and IMU data. Processing this video with VIO SLAM provides high-quality tracking of each egocentric camera through space. The first-person view from one individual provides third-person observations of other people, although these exocentric observations are sparse, intermittent, and of highly variable reliability as both cameras and subjects move. To integrate these synchronized data streams, we propose a diffusion-based approach that fuses estimates of pose based on head motion derived from egocentric camera motion with exocentric pose observations, conditioning on both observation content and reliability. Our model is trained on a mixture of single-person motion-capture data and multi-person video in order to learn rich priors for body motion trajectories and video observation reliability. Evaluation on challenging multi-person datasets suggests our fusion approach improves over motion-only and vision-only baselines in terms of both absolute and relative pose accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。