用强化学习让机器人在变化水流中高效导航,发现感知速度省力,感知涡度更准。
Flow-aware Optimal Navigation in Unsteady Flows through Reinforcement Learning

- 用TD3算法训练智能体,仅靠局部流速或涡度信息导航
- 感知流速的智能体最节能,感知涡度的最接近目标
- 给全局流参数反而降低性能,暗示隐式表征更鲁棒
自主机器人在非平稳时变流体中的导航仍是根本性挑战,源于部分可观测性与真实环境的不可预测性。传统最优控制需事先知晓全局流场,而生物系统则依赖局部感官线索成功导航。本文提出基于TD3算法的强化学习方法,训练智能体在参数化混沌双涡流场中抵达任意目标。评估五种类生物流感策略(相对位置、局部速度、局部涡度及短期记忆变体),并分析提供全局流参数的影响。数值结果表明,能感知并记忆一定数量流速测量值的智能体表现最佳。实验揭示传感器效用权衡:速度感知者优化能耗,涡度感知者具更优结构映射能力且逼近目标更佳。提供显式全局流参数反而降低导航性能,提示强化学习系统在受限于隐式流表征时可学习更具鲁棒性与泛化性的策略。研究为生物启发式机器人导航从仿真向现实环境过渡提供洞见。
原文摘要 · Abstract (English)
Autonomous robotic navigation in nonstationary time-varying fluid flows remains a fundamental challenge due to partial observability and the unpredictability of realistic environments. While classical optimal control frameworks employed in robotics require unrealistic a-priori global flow knowledge, biological systems are able to navigate successfully by exploiting localized sensory cues. In this work we present a reinforcement learning approach using the TD3 algorithm to train autonomous agents to reach arbitrary targets within a parametric, chaotic double-gyre flow. To investigate optimal sensory mechanisms, we evaluate five bio-inspired observation strategies based on relative position, local velocity or local vorticity measures, and short-term memory variants. Additionally, we analyze the impact of providing agents with explicit global flow parameters. Numerical results demonstrate that an agent that is able to sense and remember a set number of flow velocity measures achieves the highest performance. The experiments reveal a trade-off in sensor utility: velocity-aware agents optimize energy efficiency, whereas vorticity sensors provide superior structural mapping and achieve better target proximity. Incorporating explicit global flow parameters is shown to decrease navigation performance. This behavior suggests that reinforcement learning-based autonomous systems develop more robust and general policies when restricted to implicit flow representations. The presented results offer insights for improving the transition of bio-inspired robotic navigation from simulation to real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。