用强化学习让机器人学会摔倒时自我保护,降低损伤风险。
Discovering Self-Protective Falling Policy for Humanoid Robot via Deep Reinforcement Learning
- 设计奖励函数与课程训练,让机器人自主学习摔倒防护策略。
- 形成三角结构可显著降低摔倒时的硬件损伤,实测效果优于传统方法。
- 方法适配真实机器人平台,适用于复杂环境下的安全落地。
近年来,类人机器人受到广泛关注并取得显著进展。然而,由于其复杂的形态、动力学特性及控制策略局限,相比四足或轮式机器人更易跌倒。其高重量、高重心和高自由度在失控跌倒时极易造成自身及周围物体严重损坏。现有研究多依赖基于控制的方法,难以应对多样化的跌倒场景,且可能引入不合适的先验人类行为。本文利用大规模深度强化学习与课程学习,设计精心的奖励函数与领域多样化训练流程,成功训练类人机器人探索并发现一种自保护的跌倒行为:通过形成“三角”结构,能显著降低刚性躯体在跌倒时的损伤。通过全面指标与实验验证,量化了性能表现,可视化了跌倒行为,并成功将该策略迁移到真实机器人平台。
原文摘要 · Abstract (English)
Humanoid robots have received significant research interests and advancements in recent years. Despite many successes, due to their morphology, dynamics and limitation of control policy, humanoid robots are prone to fall as compared to other embodiments like quadruped or wheeled robots. And its large weight, tall Center of Mass, high Degree-of-Freedom would cause serious hardware damages when falling uncontrolled, to both itself and surrounding objects. Existing researches in this field mostly focus on using control based methods that struggle to cater diverse falling scenarios and may introduce unsuitable human prior. On the other hand, large-scale Deep Reinforcement Learning and Curriculum Learning could be employed to incentivize humanoid agent discovering falling protection policy that fits its own nature and property. In this work, with carefully designed reward functions and domain diversification curriculum, we successfully train humanoid agent to explore falling protection behaviors and discover that by forming a `triangle' structure, the falling damages could be significantly reduced with its rigid-material body. With comprehensive metrics and experiments, we quantify its performance with comparison to other methods, visualize its falling behaviors and successfully transfer it to real world platform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。