用惯性传感器同步捕捉人与物体交互的运动,解决遮挡问题。
IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion

- 从惯性数据直接推断手物接触,用接触信号指导动作推理
- 三阶段融合提升人体姿态与物体轨迹精度,抑制漂移
- 可无缝接入现有系统,适合需要真实交互捕捉的应用
在增强现实、虚拟现实和机器人应用中,捕捉带物体交互的全身运动至关重要,但传统视觉方法因遮挡和空间限制难以实现。惯性测量单元(IMUs)无需视线即可工作,但现有方法仅处理孤立人体,忽略物体接触与动态。为此,我们提出 IMU-HOI 框架,通过人体与物体上的稀疏 IMU,联合恢复全身姿态与 6-DoF 物体轨迹,并显式建模人-物交互。该方法首先从 IMU 流中推断手物接触的概率,作为高层信号引导运动推理路径;再通过三阶段融合流程,优化人体姿态与根部位移,并融合手部前向运动学与物体 IMU 数据,获得一致且抗漂移的主客体运动轨迹。在复杂人-物交互场景下的实验表明,相比已有惯性捕获方法,精度显著提升。此外,IMU-HOI 可以以最小改动嵌入现有稀疏 IMU 动作捕捉系统,将纯惯性捕捉的范围从孤立人体扩展至完整的人-物交互与联合运动估计。
原文摘要 · Abstract (English)
Capturing full-body human motion with object interactions is crucial for AR/VR and robotics applications, yet it remains challenging for conventional vision-based methods due to occlusions and constrained capture volumes. Inertial measurement units (IMUs) offer a compelling alternative without line-of-sight requirements, but existing IMU-based motion capture assumes an isolated human and ignores object contacts and dynamics. To bridge this gap, we present IMU-HOI, a novel framework that jointly recovers full-body human pose and 6-DoF object trajectory from sparse IMUs on the body and object, explicitly modeling human-object interaction. Our approach first infers probabilistic hand-object contacts directly from IMU streams and uses them as a high-level signal to route between kinematic and inertial reasoning. These contact cues drive a three-stage fusion pipeline that refines human pose and root translation, and fuses hand-based forward kinematics with object-IMU integration for object motion, yielding coherent, drift-resilient trajectories for both human and object. Experiments on challenging human-object interaction scenarios demonstrate substantial accuracy gains over prior inertial motion capture methods. Moreover, IMU-HOI can be plugged into existing sparse-IMU mocap backbones with minimal changes, effectively extending the scope of purely inertial motion capture from isolated humans to full human-object interaction and joint motion estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。