用压力信号提升人形机器人动作模仿的物理一致性。
PressMimic: Pressure-Guided Motion Capture and Control for Humanoid Robot Imitation

- 融合视觉与压力数据,提升动作捕捉精度与接触判断。
- 压力监督强化学习策略,减少足部滑动和穿透等异常行为。
- 适合关注机器人物理交互与真实动作还原的研究者。
人形机器人动作模仿不仅需要准确感知人体运动学,还需忠实再现与环境的物理交互。现有方法主要依赖视觉动作捕捉与运动学模仿,忽视接触动力学,导致足部滑动、穿地及不稳定等伪影。本文从物理基础出发,引入压力作为感知与控制的统一模态。提出PressMimic框架,将压力贯穿于从动作捕捉到机器人控制的全流程。感知阶段,设计FRAPPE++多模态模型,融合RGB与压力数据,联合估计3D姿态与全局运动,压力提供明确的接触与支撑约束,解决纯视觉估计中的歧义。控制阶段,提出压力监督策略(PSP),将压力信号融入强化学习,实现执行过程中的物理一致接触模式。进一步构建MotionPRO大规模数据集,包含同步的RGB、压力与动作捕捉数据。实验表明,压力显著提升动作估计精度、轨迹一致性与执行稳定性。结果证明,压力是有效的物理锚定信号,可贯通感知与控制,实现物理一致的人形动作模仿。
原文摘要 · Abstract (English)
Humanoid motion imitation requires not only accurate perception of human kinematics but also faithful reproduction of physical interactions with the environment. However, existing pipelines rely primarily on vision-based motion capture and kinematic imitation, largely ignoring contact dynamics, leading to artifacts such as foot sliding, floor penetration, and unstable behaviors. In this work, we revisit humanoid motion imitation from the perspective of physical grounding and leverage pressure as a unified modality across perception and control. We present PressMimic, a framework that integrates pressure into the full pipeline from motion capture to humanoid control. In the perception stage, we introduce FRAPPE++, a multimodal model that fuses RGB and pressure to jointly estimate 3D pose and global motion, where pressure provides explicit contact and support constraints to resolve ambiguity in vision-based estimation. In the control stage, we propose a pressure-supervised policy (PSP) that incorporates pressure-derived signals into reinforcement learning, enabling physically consistent contact patterns during execution. We further construct MotionPRO, a large-scale dataset with synchronized RGB, pressure, and motion capture data. Experiments show that pressure improves motion estimation accuracy, trajectory consistency, and execution stability. These results demonstrate that pressure serves as an effective physical grounding signal, bridging perception and control for physically consistent humanoid motion imitation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。