arXiv:2509.20095cs.AI2025-09被引 2

用强化学习解释线虫群集的化学信号,发现探索者能提升群体适应力。

From Pheromones to Policies: Reinforcement Learning for Engineered Biological Swarms

  • 将化学信号视为分布式奖励,建模为强化学习中的交叉更新机制。
  • 静态环境下模型精准复现线虫觅食行为,动态中信号会固化错误选择。
  • 引入少数不敏感探索者可恢复群体灵活性,适合生物机器人设计参考。

群体智能源于简单个体间的去中心化交互,实现集体问题求解。本研究建立了线虫(C. elegans)中依赖信息素的聚集行为与强化学习(RL)之间的理论等价性,证明了刻板信号可作为分布式奖励机制。我们构建了执行觅食任务的工程化线虫群模型,发现信息素动态在数学上等同于强化学习的核心算法——交叉学习更新。基于文献数据的实验验证表明,该模型在静态条件下能准确复现真实的C. elegans觅食模式。在动态环境中,持续的信息素轨迹形成正反馈回路,导致群体固守过时策略而难以适应。通过多臂赌博机场景的计算实验发现,引入少量对信息素不敏感的探索型个体,可恢复群体的集体可塑性,实现快速任务切换。这种行为异质性平衡了探索与利用的权衡,实现了群体层面过时策略的自然淘汰。结果表明,刻板系统本质上编码了分布式强化学习过程,环境信号充当集体信用分配的外部记忆。该工作连接合成生物学与群体机器人学,推动了在动荡环境中具备鲁棒决策能力的可编程生命系统的发展。

原文摘要 · Abstract (English)

Swarm intelligence emerges from decentralised interactions among simple agents, enabling collective problem-solving. This study establishes a theoretical equivalence between pheromone-mediated aggregation in \celeg\ and reinforcement learning (RL), demonstrating how stigmergic signals function as distributed reward mechanisms. We model engineered nematode swarms performing foraging tasks, showing that pheromone dynamics mathematically mirror cross-learning updates, a fundamental RL algorithm. Experimental validation with data from literature confirms that our model accurately replicates empirical \celeg\ foraging patterns under static conditions. In dynamic environments, persistent pheromone trails create positive feedback loops that hinder adaptation by locking swarms into obsolete choices. Through computational experiments in multi-armed bandit scenarios, we reveal that introducing a minority of exploratory agents insensitive to pheromones restores collective plasticity, enabling rapid task switching. This behavioural heterogeneity balances exploration-exploitation trade-offs, implementing swarm-level extinction of outdated strategies. Our results demonstrate that stigmergic systems inherently encode distributed RL processes, where environmental signals act as external memory for collective credit assignment. By bridging synthetic biology with swarm robotics, this work advances programmable living systems capable of resilient decision-making in volatile environments.

强化学习群体智能生物机器人信息素

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。