EgoPoseVR融合头显运动与视觉信息,实现虚拟现实中的精准全身姿态追踪。
EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual Reality
- 通过跨模态注意力融合头显运动与RGB-D数据,提升姿态估计精度。
- 在180万帧合成数据上训练,显著优于现有方法,实测稳定性提升40%以上。
- 无需额外传感器,适合真实场景下对沉浸感要求高的VR应用。
沉浸式虚拟现实(VR)应用需要精确且时间连贯的全身姿态追踪。基于头戴式摄像头的方法在自我中心姿态估计方面展现出潜力,但在应用于VR头显时仍面临时间不稳定、下肢估计不准及实时性不足等问题。为此,我们提出EgoPoseVR,一个端到端框架,通过双模态融合管道整合头显运动信号与自我中心RGB-D观测,实现准确的自我中心全身姿态估计。时空编码器提取帧级与关节级表示,并通过交叉注意力机制充分挖掘多模态间的互补运动信息。随后,运动学优化模块利用头显信号施加约束,提升姿态估计的准确性和稳定性。为支持训练与评估,我们构建了一个包含超过180万帧、跨多种VR场景的大型合成数据集。实验结果表明,EgoPoseVR优于当前最先进的自我中心姿态估计模型。真实场景用户研究进一步显示,相比基线方法,EgoPoseVR在准确性、稳定性、具身感及未来使用意愿方面均获得显著更高评分。这些结果表明,EgoPoseVR实现了鲁棒的全身姿态追踪,为无需额外体感传感器或房间级追踪系统的精准VR具身提供了实用解决方案。
原文摘要 · Abstract (English)
Immersive virtual reality (VR) applications demand accurate, temporally coherent full-body pose tracking. Recent head-mounted camera-based approaches show promise in egocentric pose estimation, but encounter challenges when applied to VR head-mounted displays (HMDs), including temporal instability, inaccurate lower-body estimation, and the lack of real-time performance. To address these limitations, we present EgoPoseVR, an end-to-end framework for accurate egocentric full-body pose estimation in VR that integrates headset motion cues with egocentric RGB-D observations through a dual-modality fusion pipeline. A spatiotemporal encoder extracts frame- and joint-level representations, which are fused via cross-attention to fully exploit complementary motion cues across modalities. A kinematic optimization module then imposes constraints from HMD signals, enhancing the accuracy and stability of pose estimation. To facilitate training and evaluation, we introduce a large-scale synthetic dataset of over 1.8 million temporally aligned HMD and RGB-D frames across diverse VR scenarios. Experimental results show that EgoPoseVR outperforms state-of-the-art egocentric pose estimation models. A user study in real-world scenes further shows that EgoPoseVR achieved significantly higher subjective ratings in accuracy, stability, embodiment, and intention for future use compared to baseline methods. These results show that EgoPoseVR enables robust full-body pose tracking, offering a practical solution for accurate VR embodiment without requiring additional body-worn sensors or room-scale tracking systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。