arXiv:2603.04073cs.RO2026-03

让四足机器人安全游泳,抑制不稳定力并提升推进效率。

Swimming Under Constraints: A Safe Reinforcement Learning Framework for Quadrupedal Bio-Inspired Propulsion

  • 用带PID调节的拉格朗日乘子约束不稳定的流体力。
  • 在真实水池实验中实现更快收敛与更高推进效率。
  • 适合研究仿生机器人在复杂流体中自主控制的学者。

仿生水下推进具有高推力和高机动性,但易受升力波动等失稳力影响,且六自由度流体耦合会加剧该问题。本文将四足游泳建模为约束优化问题,旨在最大化前向推力同时最小化失稳波动。提出ACPPO-PID框架:通过PID调节的拉格朗日乘子施加约束,利用条件非对称裁剪加速学习,通过周期内几何聚合稳定更新。模型以模仿学习初始化,并经实地拖曳水池实验优化,成功迁移到自由游泳测试。结果表明,相比最先进基线,该方法显著提升推力效率、降低失稳力并加快收敛,证明了约束感知安全强化学习在复杂流体环境中实现鲁棒、可泛化的仿生运动的重要性。

原文摘要 · Abstract (English)

Bio-inspired aquatic propulsion offers high thrust and maneuverability but is prone to destabilizing forces such as lift fluctuations, which are further amplified by six-degree-of-freedom (6-DoF) fluid coupling. We formulate quadrupedal swimming as a constrained optimization problem that maximizes forward thrust while minimizing destabilizing fluctuations. Our proposed framework, Accelerated Constrained Proximal Policy Optimization with a PID-regulated Lagrange multiplier (ACPPO-PID), enforces constraints with a PID-regulated Lagrange multiplier, accelerates learning via conditional asymmetric clipping, and stabilizes updates through cycle-wise geometric aggregation. Initialized with imitation learning and refined through on-hardware towing-tank experiments, ACPPO-PID produces control policies that transfer effectively to quadrupedal free-swimming trials. Results demonstrate improved thrust efficiency, reduced destabilizing forces, and faster convergence compared with state-of-the-art baselines, underscoring the importance of constraint-aware safe RL for robust and generalizable bio-inspired locomotion in complex fluid environments.

强化学习仿生机器人水下推进安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。