用深度强化学习让蛇形机器人自适应变化的黏性环境,突破传统控制方法局限。
Adaptive Undulatory Locomotion of Snake-like Robots in Dynamic Viscous Environments via Deep Reinforcement Learning

- 通过师生架构,从模拟器中提炼环境信息,仅用本体感觉实现自适应
- 在10⁻⁷到10⁻² m²/s黏度范围内,自主生成非正弦摆动步态
- 相比传统方法提升推进速度与运输效率,适合复杂流体环境机器人设计
本文展示深度强化学习(DRL)如何使蛇形机器人在动态变化的黏性环境中实现自适应运动,克服经典预设控制方法的性能瓶颈。由于缺乏直接的流体属性传感器,该任务被建模为部分可观测马尔可夫决策过程。采用非对称演员-评论家框架,利用仅在物理仿真器中可用的特权信息训练教师策略,并将其知识蒸馏至仅依赖本体感觉信息的学生策略。在动态黏度范围(10⁻⁷至10⁻² m²/s)内的仿真结果表明,DRL智能体自主学习出非正弦的自适应步态,显著提升推进速度与运输效率,突破了传统正弦及运动学控制的固有限制。研究证实,通过特权信息蒸馏实现隐式环境推断,是应对不可预测流体动力学约束的有效方法。
原文摘要 · Abstract (English)
This paper demonstrates how deep reinforcement learning (DRL) enables adaptive locomotion of snake-like robots in dynamically changing viscous environments, overcoming the inherent performance limitations of classical predefined control methods. The lack of direct onboard sensors for fluid properties necessitates formulating this task as a partially observable Markov decision process. By employing an asymmetric actor-critic framework, a teacher policy trained using privileged information available only in the physics simulator distills its knowledge into a student policy that relies solely on proprioceptive sensor information. Simulation results across a wide range of dynamic viscosity changes ($10^{-7}$ to $10^{-2} m^2/s$) reveal that the DRL agent autonomously acquires non-sinusoidal adaptive gaits. These gaits improve propulsion velocity and transport efficiency, breaking the inherent limits of conventional sinusoidal and kinematic control. The findings establish that implicit environment inference via privileged information distillation is an effective approach to bypass the constraints of classical models under unpredictable fluid dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。