用智能眼镜实现多人协同人体动作捕捉,兼顾第一人称与第三人称视角。
EgoExoMoCap: Distributed Ego-Exo Human Motion Capture

- 利用佩戴者头戴设备的追踪信号和图像特征联合建模
- 在真实场景下实现高鲁棒性动作重建,支持多人协作
- 无需专业设备,适合真实世界应用研究
从头戴设备(HMD)中进行人体动作捕捉为具身智能和虚拟/增强现实应用提供了获取真实世界人类动作与交互数据的可扩展方式。现有方法主要关注第一人称(egocentric)身体追踪或第三人称(exocentric)环境内他人动作捕捉,两者长期独立发展。本文提出一种新型分布式框架EgoExoMoCap,首次联合利用第一人称与第三人称多模态信号,从头戴设备中估计人体运动。该方法仅需两人各佩戴一副智能眼镜即可运行,不依赖复杂多摄像机系统或侵入式动捕服。通过结合头部(及可能的手腕)追踪信号估计三维世界中的全局运动,并利用基于DINOv3的上下文感知图像特征,在噪声和遮挡条件下实现鲁棒性。在两个真实场景数据集上的大量实验表明,该方法在复杂场景下仍能稳定重建动作。
原文摘要 · Abstract (English)
Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which is crucial for applications in embodied AI and VR/AR. Existing approaches focus on either egocentric body tracking, estimating the motion of the subject wearing the device, or exocentric tracking, capturing the movements of people in the wearer's surroundings. So far, these two paradigms have largely been explored in isolation. In this paper, we propose a novel distributed framework that jointly leverages ego- and exocentric multi-modal signals for human motion estimation from HMDs. Unlike traditional motion capture systems requiring bulky multi-camera setups or obtrusive mocap suits, our approach, EgoExoMoCap, is as simple as two (or more) people, each wearing a pair of smart glasses. The method leverages head (plus potentially wrist) tracking signals for accurate estimation of global motion in the 3D world and combines context-aware image features based on DINOv3 to achieve robustness in the presence of noise and occlusions. Extensive experiments on two in-the-wild datasets show that our approach can robustly reconstruct motion even in challenging scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。