让仿人机器人同时精准操作并主动施力,适用于高负载工业场景。
Kinematics-Aware Multi-Policy Reinforcement Learning for Force-Capable Humanoid Loco-Manipulation
- 分三阶段训练上肢、下肢与指令微调策略,解耦控制复杂度。
- 上肢策略通过启发式奖励函数加速收敛,性能更优。
- 下肢采用基于力的课程学习,实现环境交互力的主动调控。
具有类人形态的仿人机器人在工业应用中潜力巨大。然而,现有运动-操作一体化方法主要关注灵巧操作,难以满足高负载工业场景中对灵巧性与主动力交互的双重需求。为此,我们提出一种基于强化学习的框架,采用解耦的三阶段训练流程:上肢策略、下肢策略和增量命令策略。为加速上肢训练,设计了一种启发式奖励函数,通过隐式嵌入正向运动学先验知识,使策略更快收敛并取得更优性能。针对下肢,开发了基于力的课程学习策略,使机器人能够主动施加并调节与环境的交互力,提升实际工况适应能力。
原文摘要 · Abstract (English)
Humanoid robots, with their human-like morphology, hold great potential for industrial applications. However, existing loco-manipulation methods primarily focus on dexterous manipulation, falling short of the combined requirements for dexterity and proactive force interaction in high-load industrial scenarios. To bridge this gap, we propose a reinforcement learning-based framework with a decoupled three-stage training pipeline, consisting of an upper-body policy, a lower-body policy, and a delta-command policy. To accelerate upper-body training, a heuristic reward function is designed. By implicitly embedding forward kinematics priors, it enables the policy to converge faster and achieve superior performance. For the lower body, a force-based curriculum learning strategy is developed, enabling the robot to actively exert and regulate interaction forces with the environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。