用动作驱动过程统一强化学习与随机过程,为脉冲神经网络提供新框架。
Action-Driven Processes for Continuous-Time Control
- 提出动作驱动过程统一强化学习与随机过程建模。
- 最小化KL散度等价于最大熵强化学习。
- 适合研究神经动力学与智能决策的交叉领域学者。
强化学习的核心是动作——根据环境观测做出的决策。动作在随机过程建模中同样关键,因其触发状态的不连续转移,并促进复杂系统中的信息流动。本文通过动作驱动过程,统一了随机过程与强化学习的视角,并将其应用于脉冲神经网络。借鉴控制即推理的思想,我们证明:对一个适当定义的动作驱动过程,最小化策略驱动的真实分布与奖励驱动的模型分布之间的Kullback-Leibler散度,等价于最大熵强化学习。
原文摘要 · Abstract (English)
At the heart of reinforcement learning are actions -- decisions made in response to observations of the environment. Actions are equally fundamental in the modeling of stochastic processes, as they trigger discontinuous state transitions and enable the flow of information through large, complex systems. In this paper, we unify the perspectives of stochastic processes and reinforcement learning through action-driven processes, and illustrate their application to spiking neural networks. Leveraging ideas from control-as-inference, we show that minimizing the Kullback-Leibler divergence between a policy-driven true distribution and a reward-driven model distribution for a suitably defined action-driven process is equivalent to maximum entropy reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。