用日常穿戴设备实现无需校准的全身动作捕捉
Human Motion Estimation with Everyday Wearables
- 融合自拍相机与可穿戴设备惯性数据,端到端建模
- 在9小时真实场景数据上实现高精度动作估计
- 适合移动健康、XR交互等实际应用场景
基于可穿戴设备的人体动作估计对扩展现实(XR)交互等应用至关重要,但现有方法普遍存在佩戴不舒适、硬件昂贵及校准繁琐等问题。为此,我们提出EveryWear,一种完全依赖智能手机、智能手表、耳机和带前后双摄像头的智能眼镜的轻量级动作捕捉方案,无需预先校准。我们构建了Ego-Elec数据集,包含9小时真实世界数据,覆盖17种室内外环境中的56项日常活动,并以动捕系统(MoCap)提供3D真值标注。方法采用多模态师生框架,结合第一人称视角视觉信息与消费级设备惯性信号。直接在真实数据上训练,有效规避了仿真到现实的差距。实验表明,该方法优于基线模型,验证了其在实用全身体感估计中的有效性。
原文摘要 · Abstract (English)
While on-body device-based human motion estimation is crucial for applications such as XR interaction, existing methods often suffer from poor wearability, expensive hardware, and cumbersome calibration, which hinder their adoption in daily life. To address these challenges, we present EveryWear, a lightweight and practical human motion capture approach based entirely on everyday wearables: a smartphone, smartwatch, earbuds, and smart glasses equipped with one forward-facing and two downward-facing cameras, requiring no explicit calibration before use. We introduce Ego-Elec, a 9-hour real-world dataset covering 56 daily activities across 17 diverse indoor and outdoor environments, with ground-truth 3D annotations provided by the motion capture (MoCap), to facilitate robust research and benchmarking in this direction. Our approach employs a multimodal teacher-student framework that integrates visual cues from egocentric cameras with inertial signals from consumer devices. By training directly on real-world data rather than synthetic data, our model effectively eliminates the sim-to-real gap that constrains prior work. Experiments demonstrate that our method outperforms baseline models, validating its effectiveness for practical full-body motion estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。