EgoPoseFormer v2 提升了头戴设备下人体动作估计的精度与稳定性。
EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR
- 基于变压器的模型融合时序一致性和空间定位,支持关键点与参数化人体表示。
- 在 EgoBody3M 基准上准确率提升 12.2%~19.4%,时间抖动降低 22.2%~51.7%。
- 自标注系统利用海量无标签数据,适合部署于 AR/VR 等真实复杂场景。
头戴视角下的人体动作估计对 AR/VR 体验至关重要,但受限于身体覆盖范围小、频繁遮挡及标注数据稀少,仍具挑战。我们提出 EgoPoseFormer v2,通过两项关键贡献应对该问题:(1) 一种基于变换器的模型,实现时序一致且空间精准的姿势估计;(2) 一个自标注系统,使模型可利用大规模无标签数据训练。该模型完全可微,引入身份条件查询、多视角空间优化、因果时序注意力,并在恒定计算开销下支持关键点与参数化人体表示。自标注系统通过不确定性感知的半监督训练,将学习规模扩展至数千万帧无标签图像。采用教师-学生架构生成伪标签并进行不确定性蒸馏,提升模型在不同环境下的泛化能力。在 EgoBody3M 基准上,模型以 0.8 毫秒延迟运行,相比两种最先进方法分别提升 12.2% 和 19.4% 的准确率,时间抖动减少 22.2% 和 51.7%。此外,自标注系统进一步将手腕 MPJPE 降低 13.1%。
原文摘要 · Abstract (English)
Egocentric human motion estimation is essential for AR/VR experiences, yet remains challenging due to limited body coverage from the egocentric viewpoint, frequent occlusions, and scarce labeled data. We present EgoPoseFormer v2, a method that addresses these challenges through two key contributions: (1) a transformer-based model for temporally consistent and spatially grounded body pose estimation, and (2) an auto-labeling system that enables the use of large unlabeled datasets for training. Our model is fully differentiable, introduces identity-conditioned queries, multi-view spatial refinement, causal temporal attention, and supports both keypoints and parametric body representations under a constant compute budget. The auto-labeling system scales learning to tens of millions of unlabeled frames via uncertainty-aware semi-supervised training. The system follows a teacher-student schema to generate pseudo-labels and guide training with uncertainty distillation, enabling the model to generalize to different environments. On the EgoBody3M benchmark, with a 0.8 ms latency on GPU, our model outperforms two state-of-the-art methods by 12.2% and 19.4% in accuracy, and reduces temporal jitter by 22.2% and 51.7%. Furthermore, our auto-labeling system further improves the wrist MPJPE by 13.1%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。