用物理流场训练微机器人集群,实现高效、平滑的逆流前进。
Micro-Swarm Locomotion Optimization in Dynamic Flow using Multi-Objective Multi-Agent Reinforcement Learning

- 结合流体模拟与多智能体强化学习,让16个微型机器人在脉动流中自主协调。
- 达成上游推进奖励6.5-7.0,能量效率0.63-0.65,运动平滑度0.97-0.99。
- 发现三种未预设的自组织行为,适合生物导航与微流控系统研究者参考。
在真实时变流体环境中协调微机器人群体,仍是生物医学与环境应用中的重大挑战。本文提出一种混合 CFD-MO-MARL(计算流体动力学-多目标-多智能体强化学习)框架,将高保真不可压缩纳维-斯托克斯求解器与去中心化的近端策略优化相结合,学习在振荡流中群体控制策略。模拟了16个磁驱动微机器人,在2毫米通道内沿脉动动脉波形导航,同时优化上游推进、能量效率和运动平滑性。通过投影冲突梯度(PCGrad)解决目标冲突,无PCGrad时能量与平滑性奖励训练中崩溃,证明梯度冲突缓解对多目标学习稳定性至关重要。收敛策略实现推进奖励6.5-7.0,能量效率0.63-0.65,平滑度0.97-0.99,主目标优于暴力基线超过8奖励单位。训练中涌现出三种未编码于奖励函数的行为:降低峰值流速的流体节流构型、利用流动反向实现逆流的周期同步棘轮机制,以及接近目标边界时的个体化最终逼近策略。结果表明,可直接将物理真实的流体-智能体交互集成至多目标强化学习,为生物导航、环境监测与微流控系统提供可扩展的微群控制框架。
原文摘要 · Abstract (English)
Coordinating micro-robotic swarms in realistic, time-dependent fluid environments remains a major challenge for biomedical and environmental applications. We present a hybrid CFD-MO-MARL (Computational Fluid Dynamics-Multi Objective-Multi Agent Reinforcement Learning) framework that couples a high-fidelity incompressible Navier--Stokes solver with decentralized proximal policy optimization to learn swarm control policies in oscillatory flow. Sixteen magnetically actuated micro-robots were simulated to navigate a pulsatile arterial waveform within a 2 mm channel while jointly optimizing upstream progression, energy efficiency, and motion smoothness. Conflicting objectives are resolved using Projected Conflicting Gradient (PCGrad) surgery. Without PCGrad, energy and smoothness rewards collapse during training, demonstrating that gradient conflict resolution is essential for stable multi-objective learning. The converged policy achieves progress rewards of 6.5-7.0, energy efficiency of 0.63-0.65, and smoothness of 0.97-0.99, outperforming brute-force baselines by more than 8 reward units on the primary objective. Training reveals three emergent behaviors not encoded in the reward function: hydrodynamic throttling formations that reduce peak flow velocities, a cycle-synchronized ratchet mechanism that exploits flow reversals for upstream movement, and individualized final-approach strategies near the target boundary. These results demonstrate that physically realistic fluid--agent interactions can be integrated directly into multi-objective reinforcement learning, providing a scalable framework for micro-swarm control in biomedical navigation, environmental monitoring, and microfluidic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。