让文字指令生成的机器人动作既符合物理规律又能准确执行。
RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
- 用物理仿真评估动作可行性,通过强化学习优化生成结果。
- 在仿真中生成可执行动作,真实机器人成功部署。
- 适合需要精准动作控制的具身智能研究者。
本文针对机器人领域一个关键挑战:将文本驱动的人体运动转化为类人机器人可执行的动作,实现新行为的高效、低成本学习。现有文本到动作生成方法虽能实现语言与动作的语义对齐,但常生成在运动学或物理上不可行的动作,难以用于真实部署。为弥合模拟到现实的差距,我们提出基于物理反馈的强化学习(RLPF)框架,将物理感知的动作评估与文本条件生成相结合。RLPF采用动作追踪策略在物理仿真环境中评估动作可行性,并生成奖励信号以微调运动生成器;同时引入对齐验证模块,确保生成动作保持与文本指令的语义一致性。联合优化保证了动作的物理合理性与指令对齐性。大量实验表明,RLPF显著优于基线方法,在保持文本对应关系的同时大幅提升了动作的物理可行性,成功实现在真实类人机器人上的部署。
原文摘要 · Abstract (English)
This paper focuses on a critical challenge in robotics: translating text-driven human motions into executable actions for humanoid robots, enabling efficient and cost-effective learning of new behaviors. While existing text-to-motion generation methods achieve semantic alignment between language and motion, they often produce kinematically or physically infeasible motions unsuitable for real-world deployment. To bridge this sim-to-real gap, we propose Reinforcement Learning from Physical Feedback (RLPF), a novel framework that integrates physics-aware motion evaluation with text-conditioned motion generation. RLPF employs a motion tracking policy to assess feasibility in a physics simulator, generating rewards for fine-tuning the motion generator. Furthermore, RLPF introduces an alignment verification module to preserve semantic fidelity to text instructions. This joint optimization ensures both physical plausibility and instruction alignment. Extensive experiments show that RLPF greatly outperforms baseline methods in generating physically feasible motions while maintaining semantic correspondence with text instruction, enabling successful deployment on real humanoid robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。