arXiv:2509.26082cs.RO2025-09被引 2

用进化强化学习让人形机器人自动优化硬件和控制,提升引体向上性能。

Evolutionary Continuous Adaptive RL-Powered Co-Design for Humanoid Chin-Up Performance

  • 结合进化算法与连续自适应强化学习,同步优化硬件与控制策略。
  • 在动态引体向上任务中实现更高得分,探索更广的设计空间。
  • 适合研究机器人协同设计、强化学习控制的科研人员。

人形机器人在设计与控制方面取得显著进展,越来越多研究强调二者协同优化以提升整体性能。传统方法采用串行流程,先完成硬件设计再开发控制算法,容易局限机器人潜力。近期方法提倡协同设计,同时优化硬件与控制以最大化性能。本文提出进化连续自适应强化学习协同设计框架(EA-CoRL),融合强化学习与进化策略,实现控制策略对硬件的持续适应。该框架包含两个核心组件:设计演化(Design Evolution),通过进化算法探索硬件配置以发现高效结构;策略连续适配(Policy Continuous Adaptation),在不断演化的设计中微调特定任务的控制策略以最大化奖励。我们在RH5人形机器人上评估了该框架,联合优化执行器(齿轮比)与控制策略,完成此前因执行器限制无法实现的高动态引体向上任务。相比现有先进强化学习协同设计方法,EA-CoRL展现出更高的适应度得分和更广泛的设计空间探索能力,凸显连续策略适配在机器人协同设计中的关键作用。

原文摘要 · Abstract (English)

Humanoid robots have seen significant advancements in both design and control, with a growing emphasis on integrating these aspects to enhance overall performance. Traditionally, robot design has followed a sequential process, where control algorithms are developed after the hardware is finalized. However, this can be myopic and prevent robots to fully exploit their hardware capabilities. Recent approaches advocate for co-design, optimizing both design and control in parallel to maximize robotic capabilities. This paper presents the Evolutionary Continuous Adaptive RL-based Co-Design (EA-CoRL) framework, which combines reinforcement learning (RL) with evolutionary strategies to enable continuous adaptation of the control policy to the hardware. EA-CoRL comprises two key components: Design Evolution, which explores the hardware choices using an evolutionary algorithm to identify efficient configurations, and Policy Continuous Adaptation, which fine-tunes a task-specific control policy across evolving designs to maximize performance rewards. We evaluate EA-CoRL by co-designing the actuators (gear ratios) and control policy of the RH5 humanoid for a highly dynamic chin-up task, previously unfeasible due to actuator limitations. Comparative results against state-of-the-art RL-based co-design methods show that EA-CoRL achieves higher fitness score and broader design space exploration, highlighting the critical role of continuous policy adaptation in robot co-design.

协同设计强化学习人形机器人进化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。