用双目相机和6个惯性传感器实现高精度、考虑人体差异的实时动作捕捉。
Stereo-Inertial Poser: Towards Metric-Accurate Shape-Aware Motion Capture Using Sparse IMUs and a Single Stereo Camera
- 双目视觉+惯性传感器融合,解决单目深度模糊问题。
- 实现200+帧/秒实时运行,长期录制无漂移,足部滑动现象减少。
- 能根据个体体型差异动态调整,适合个性化动作建模场景。
近年来视觉-惯性动作捕捉系统利用单目相机与稀疏惯性测量单元(IMUs)组合,成为低成本解决方案,有效缓解单一模态系统的遮挡与漂移问题。然而,仍受限于单目深度模糊导致的全局平移度量不准,以及忽略人体解剖学差异的局部运动估计。我们提出 Stereo-Inertial Poser,一个基于单个双目相机和六个IMU的实时动作捕捉系统,可实现度量准确且考虑体型差异的3D人体动作估计。通过将单目RGB替换为双目视觉,系统借助校准基线几何解决深度模糊,实现直接3D关键点提取与身体形状参数估计。惯性数据与视觉信息融合用于预测补偿漂移的关节位置与根节点运动,同时引入新颖的体型感知融合模块,动态调和解剖学差异与全局平移。端到端流程无需优化后处理,实测超过200 FPS,支持实时部署。多数据集定量评估显示其性能达到当前最优水平。定性结果表明,长时间录制下无漂移,且显著减少足部滑动现象。
原文摘要 · Abstract (English)
Recent advancements in visual-inertial motion capture systems have demonstrated the potential of combining monocular cameras with sparse inertial measurement units (IMUs) as cost-effective solutions, which effectively mitigate occlusion and drift issues inherent in single-modality systems. However, they are still limited by metric inaccuracies in global translations stemming from monocular depth ambiguity, and shape-agnostic local motion estimations that ignore anthropometric variations. We present Stereo-Inertial Poser, a real-time motion capture system that leverages a single stereo camera and six IMUs to estimate metric-accurate and shape-aware 3D human motion. By replacing the monocular RGB with stereo vision, our system resolves depth ambiguity through calibrated baseline geometry, enabling direct 3D keypoint extraction and body shape parameter estimation. IMU data and visual cues are fused for predicting drift-compensated joint positions and root movements, while a novel shape-aware fusion module dynamically harmonizes anthropometry variations with global translations. Our end-to-end pipeline achieves over 200 FPS without optimization-based post-processing, enabling real-time deployment. Quantitative evaluations across various datasets demonstrate state-of-the-art performance. Qualitative results show our method produces drift-free global translation under a long recording time and reduces foot-skating effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。