arXiv:2605.15938physics.bio-phcs.LG2026-05

用简单记忆的强化学习实现湍流中嗅觉寻源,模仿昆虫行为。

Clock-state olfactory search in turbulent flows using Q-learning: The geometry of plume recovery

论文配图:Clock-state olfactory search in turbulent flows using Q-learning: The geometry of plume recovery
图 1 · 摘自论文原文
  • 仅记录上次闻到气味的时间,用Q-learning学习导航策略。
  • 在湍流模拟中成功恢复气味羽流,表现接近昆虫本能行为。
  • 方法简洁可解释,适合研究生物启发式搜索与机器人寻源。

在湍流中寻找气味源需有效利用嗅觉观测的历史信息以形成稳健的导航策略。本文使用表格型Q-learning训练一个嗅觉搜索智能体,其仅保留上次闻到气味以来的运行时钟作为最小记忆。该智能体学习到一种可解释的羽流恢复策略,结合了昆虫已知的行为模式:上行、横向搜索和顺风返回。尽管在直接数值模拟生成的湍流数据中表现良好,但智能体因无法根据局部间歇性水平自适应调整策略而受限;我们证明增加策略灵活性可显著提升鲁棒性。

原文摘要 · Abstract (English)

Finding an odor source in a turbulent flow requires effectively leveraging the history of olfactory observations into a robust navigation strategy. In this work, we use tabular Q-learning to train an olfactory search agent with a minimal memory of past observations: only a running clock since the last whiff. This agent learns an interpretable strategy to recover the plume which combines well-known behaviors observed in insects: surging, casting, and a return downwind. While achieving good performance on data from direct numerical simulations of turbulence, the agent is limited by an inability to adapt its strategy to the local intermittency level; we show that providing more flexibility improves robustness.

强化学习嗅觉寻源湍流导航生物启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。