用生化反应网络模拟藻类光趋性,揭示游动-翻滚是主动探知环境的策略。
Implementation of reinforcement learning in chemical reaction networks: application to phototaxis as curiosity-driven exploration

- 将光趋性建模为信息驱动的感知-动作过程,通过生化反应网络实现
- 基于30条实验轨迹反推行为目标,重现真实光照对齐分布
- 首次展示细胞内生化网络可支持主动探索与信息获取
生命系统在噪声和不完整感官信号中导航。单细胞绿藻的光趋性常被建模为由刺激-响应规则驱动的运行-翻滚过程,但这类描述忽略了生物体主动采样环境以减少感知模糊的能力。从最小认知视角出发,我们将此导航重新理解为一种主观的信息驱动感知-动作过程。为此,我们提出一个将部分可观测马尔可夫决策过程(POMDP)与生化反应动力学相连接的框架:环境变量不可见,细胞通过无记忆贝叶斯更新机制从每次观测中构建最小内部状态。该内部动态平衡向光定向与探索性重定向,可通过化学反应网络常微分方程(CRN-ODEs)实现。模型包含生物物理的光感受过程及化学可计算的信息增益多项式上界。利用逆强化学习(IRL)分析30条实验记录的衣藻轨迹,我们推断出与观察到的光趋性运动一致的行为目标,并与标准随机模拟算法(SSA)基线进行对比。结果表明,本模型再现了经验性的光照对齐分布,性能与目标SSA基线相当。在此框架下,运行-翻滚交替表现为一种信息获取策略:翻滚重新定向细胞,以采样新的感官配置并解决传感器模糊问题,展示了细胞内生化网络如何支持适应性信息寻求行为。
原文摘要 · Abstract (English)
Living systems navigate environments using noisy and incomplete sensory signals. In unicellular algae, phototaxis is often modeled as a mechanistic run--tumble process driven by stimulus--response rules. However, such descriptions overlook how organisms actively sample their environment to reduce sensory ambiguity. From a minimal cognition perspective, we reframe this navigation as a subjective, information-driven sensorimotor process. To this end, we propose a framework linking a Partially Observable Markov Decision Process (POMDP) with biochemical reaction dynamics. Environmental variables are hidden, while the cell updates a minimal internal state from each observation through a memoryless Bayesian step. These internal dynamics balance orienting toward light with exploratory reorientation and can be implemented through Chemical-Reaction-Network Ordinary Differential Equations (CRN--ODEs). Our model includes a biophysical observation process for photoreception and a chemically computable polynomial bound on information gain. Using Inverse Reinforcement Learning (IRL) on 30 experimentally recorded Chlamydomonas trajectories, we infer the behavioral objective consistent with observed phototactic motion and benchmark the resulting dynamics with standard Stochastic Simulation Algorithm (SSA) baselines. Our model reproduces the empirical alignment-to-light distribution, comparable to objective SSA baselines on this dataset. Within this framework, run--tumble alternation emerges as an information-acquisition strategy: tumbling reorients the cell to sample new sensory configurations and resolve sensor ambiguity, demonstrating how intracellular biochemical networks can support adaptive information-seeking behavior in cellular navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。