用可穿戴设备融合传感器,实现野外人类动作的高精度全身三维姿态捕捉。
RoSHI: A Versatile Robot-oriented Suit for Human Data In-the-Wild
- 融合低成本惯性传感器与眼镜摄像头,实现全局坐标下的全身姿态估计。
- 在敏捷动作数据集上性能优于其他眼动基准,接近顶尖外置系统表现。
- 采集数据可用于真实人形机器人策略学习,适合机器人动作数据收集场景。
扩大机器人学习规模可能需要包含丰富且长时程自然交互的人类数据。现有数据采集方法在便携性、遮挡鲁棒性和全局一致性之间存在权衡。我们提出RoSHI,一种混合式可穿戴装置,将低成本稀疏惯性测量单元(IMUs)与Project Aria眼镜融合,通过第一人称视角感知,估计佩戴者在度量全局坐标系下的完整3D姿态和身体形状。该设计基于两类传感器的互补性:IMUs提供对遮挡和高速运动的鲁棒性,而第一人称SLAM则锚定长时间运动并稳定上半身姿态。我们构建了一个包含敏捷动作的基准数据集以评估RoSHI。在该数据集上,我们的方法普遍优于其他第一人称基线,性能接近最先进的外置基线(SAM3D)。最后,我们验证了该系统记录的动作数据适用于真实世界人形机器人策略学习。更多视频、数据请访问项目网页:https://roshi-mocap.github.io/
原文摘要 · Abstract (English)
Scaling up robot learning will likely require human data containing rich and long-horizon interactions in the wild. Existing approaches for collecting such data trade off portability, robustness to occlusion, and global consistency. We introduce RoSHI, a hybrid wearable that fuses low-cost sparse IMUs with the Project Aria glasses to estimate the full 3D pose and body shape of the wearer in a metric global coordinate frame from egocentric perception. This system is motivated by the complementarity of the two sensors: IMUs provide robustness to occlusions and high-speed motions, while egocentric SLAM anchors long-horizon motion and stabilizes upper body pose. We collect a dataset of agile activities to evaluate RoSHI. On this dataset, we generally outperform other egocentric baselines and perform comparably to a state-of-the-art exocentric baseline (SAM3D). Finally, we demonstrate that the motion data recorded from our system are suitable for real-world humanoid policy learning. For videos, data and more, visit the project webpage: https://roshi-mocap.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。