arXiv:2501.03687q-bio.CBcs.LG2025-01被引 4

用强化学习模拟细菌趋化运动,优化导航策略。

Run-and-tumble chemotaxis using reinforcement learning

  • 设计基于轨迹历史的奖励机制,控制前进与转向动作。
  • 在不同梯度环境下,找到探索与利用的最佳平衡点。
  • 适用于需要高效环境感知的移动智能体研究。

细菌通过‘跑动-翻滚’运动模式沿吸引剂浓度梯度上行:延长顺梯度运行时间、缩短逆梯度运行时间,实现向高浓度区域迁移。受此启发,我们构建了一个一维强化学习(RL)框架,其中智能体在吸引剂梯度环境中移动,可执行两种动作:保持方向持续前进或反转方向。根据智能体轨迹的近期历史为动作分配代价。研究问题为:在不同吸引剂分布下,哪种RL策略表现最优?评估指标包括:(a) 长时间后能否定位到有利区域,(b) 能否充分学习环境全貌。结果表明,需根据吸引剂分布与初始条件,在探索与利用之间取得最优权衡,以实现最高效性能。

原文摘要 · Abstract (English)

Bacterial cells use run-and-tumble motion to climb up attractant concentration gradient in their environment. By extending the uphill runs and shortening the downhill runs the cells migrate towards the higher attractant zones. Motivated by this, we formulate a reinforcement learning (RL) algorithm where an agent moves in one dimension in the presence of an attractant gradient. The agent can perform two actions: either persistent motion in the same direction or reversal of direction. We assign costs for these actions based on the recent history of the agent's trajectory. We ask the question: which RL strategy works best in different types of attractant environment. We quantify efficiency of the RL strategy by the ability of the agent (a) to localize in the favorable zones after large times, and (b) to learn about its complete environment. Depending on the attractant profile and the initial condition, we find an optimum balance is needed between exploration and exploitation to ensure the most efficient performance.

强化学习趋化运动智能体导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。