arXiv:2509.19525cs.RO2025-09被引 3

用实时强化学习让软机器人在单次部署中自适应动态平衡。

Real-Time Reinforcement Learning for Dynamic Tasks with a Parallel Soft Robot

  • 通过课程学习扩展平衡点邻域,实现单次硬件部署即学即用。
  • 15分钟内训练完成,即使一半驱动器失效仍能稳定平衡。
  • 适合需快速适应复杂环境的柔性机器人系统研发者。

闭环控制仍是软体机器人领域的开放挑战。软致动器在动态载荷下的非线性响应限制了解析模型的应用。传统控制方法为规避非线性、滞后、大变形及损伤风险,过度简化了构型空间利用。此外,传统的基于数据的强化学习(RL)方法受限于样本效率和初始化不一致。本文展示了一种实时强化学习方法,可在单次硬件部署中可靠学习动态平衡任务的控制策略。我们采用基于电机驱动手性剪切辅助结构(HSA)的3D打印平行软致动器构建可变形斯特林平台。通过引入基于已知平衡点邻域扩展的课程学习策略,实现了任意坐标下的可靠单次部署平衡。除了对比基于模型与无模型方法的表现,我们还证明:在单次部署中,最大扩散强化学习(Maximum Diffusion RL)可在半数致动器失效(通过屈曲或剪断螺栓)后仍成功学习动态平衡,训练无需先验数据,最快仅需15分钟,性能几乎等同于完整平台。该单次学习方案使软体机器人能在真实世界中可靠学习,推动更丰富、更强大的软体机器人发展。

原文摘要 · Abstract (English)

Closed-loop control remains an open challenge in soft robotics. The nonlinear responses of soft actuators under dynamic loading conditions limit the use of analytic models for soft robot control. Traditional methods of controlling soft robots underutilize their configuration spaces to avoid nonlinearity, hysteresis, large deformations, and the risk of actuator damage. Furthermore, episodic data-driven control approaches such as reinforcement learning (RL) are traditionally limited by sample efficiency and inconsistency across initializations. In this work, we demonstrate RL for reliably learning control policies for dynamic balancing tasks in real-time single-shot hardware deployments. We use a deformable Stewart platform constructed using parallel, 3D-printed soft actuators based on motorized handed shearing auxetic (HSA) structures. By introducing a curriculum learning approach based on expanding neighborhoods of a known equilibrium, we achieve reliable single-deployment balancing at arbitrary coordinates. In addition to benchmarking the performance of model-based and model-free methods, we demonstrate that in a single deployment, Maximum Diffusion RL is capable of learning dynamic balancing after half of the actuators are effectively disabled, by inducing buckling and by breaking actuators with bolt cutters. Training occurs with no prior data, in as fast as 15 minutes, with performance nearly identical to the fully-intact platform. Single-shot learning on hardware facilitates soft robotic systems reliably learning in the real world and will enable more diverse and capable soft robots.

强化学习软体机器人实时控制自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。