用压力信号提升动作捕捉的物理合理性,让虚拟人和机器人更真实稳定。
MotionPRO: Exploring the Role of Pressure in Human MoCap and Beyond
- 通过压力、视觉与光学数据构建大规模动作捕捉数据集MotionPRO。
- 仅用压力信号即可准确估计全身姿态与全局轨迹,融合压力与视觉可显著提升精度。
- 适合研究虚拟人驱动、机器人控制及具身智能的开发者使用。
现有动作捕捉方法多关注视觉相似性,忽视物理合理性,导致虚拟人或人形机器人在三维场景中出现时间漂移、抖动、滑动、穿透等问题,且全局轨迹不准确。本文从人体与物理世界交互角度出发,探索压力信号的作用。首先,构建包含70名志愿者、400种动作、总计1240万帧姿态的大规模带压力、RGB和光学传感器的动作捕捉数据集MotionPRO。其次,通过两项挑战性任务验证压力信号的必要性与有效性:(1)仅基于压力进行姿态与轨迹估计,提出结合小卷积核解码器与长短时注意力模块的网络,证明压力能提供准确全局轨迹与合理下肢姿态;(2)融合压力与RGB信号进行估计,通过沿相机轴的正交相似性约束与沿垂直轴的全身接触约束,增强跨注意力机制以融合特征。实验表明,融合压力与RGB不仅显著提升客观指标,还能合理驱动3D场景中的虚拟人(SMPL)。此外,引入物理感知使机器人执行更精准稳定的动作,对具身人工智能发展具有重要意义。
原文摘要 · Abstract (English)
Existing human Motion Capture (MoCap) methods mostly focus on the visual similarity while neglecting the physical plausibility. As a result, downstream tasks such as driving virtual human in 3D scene or humanoid robots in real world suffer from issues such as timing drift and jitter, spatial problems like sliding and penetration, and poor global trajectory accuracy. In this paper, we revisit human MoCap from the perspective of interaction between human body and physical world by exploring the role of pressure. Firstly, we construct a large-scale human Motion capture dataset with Pressure, RGB and Optical sensors (named MotionPRO), which comprises 70 volunteers performing 400 types of motion, encompassing a total of 12.4M pose frames. Secondly, we examine both the necessity and effectiveness of the pressure signal through two challenging tasks: (1) pose and trajectory estimation based solely on pressure: We propose a network that incorporates a small kernel decoder and a long-short-term attention module, and proof that pressure could provide accurate global trajectory and plausible lower body pose. (2) pose and trajectory estimation by fusing pressure and RGB: We impose constraints on orthographic similarity along the camera axis and whole-body contact along the vertical axis to enhance the cross-attention strategy to fuse pressure and RGB feature maps. Experiments demonstrate that fusing pressure with RGB features not only significantly improves performance in terms of objective metrics, but also plausibly drives virtual humans (SMPL) in 3D scene. Furthermore, we demonstrate that incorporating physical perception enables humanoid robots to perform more precise and stable actions, which is highly beneficial for the development of embodied artificial intelligence. Project page is available at: https://nju-cite-mocaphumanoid.github.io/MotionPRO/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。