用深度循环强化学习模拟灵长类决策,发现其机制可高效应对噪声信息。
Primate-like perceptual decision making emerges through deep recurrent reinforcement learning
- 通过端到端深度循环网络在噪声感知任务中强化学习
- 网络学会速度与准确率权衡、灵活修正判断等灵长类决策能力
- 内部动态与灵长类神经生理结果相似,支持进化压力假说
尽管对灵长类决策的神经机制已有深入理解,但其演化原因仍不明确。理论认为,这些机制是为了在噪声且随时间变化的信息中最大化奖励而演化而来。为验证该理论,我们在一个噪声感知辨别任务上,使用强化学习训练了一个端到端的深度循环神经网络。训练后的网络习得了灵长类决策的关键特征:在速度与准确率之间权衡,并能根据新信息灵活修正判断。网络内部动态显示,这些能力由与灵长类神经生理研究中观察到的类似决策机制支撑。结果为塑造灵长类灵活决策能力的关键进化压力提供了实验支持。
原文摘要 · Abstract (English)
Progress has led to a detailed understanding of the neural mechanisms that underlie decision making in primates. However, less is known about why such mechanisms are present in the first place. Theory suggests that primate decision making mechanisms, and their resultant behavioral abilities, emerged to maximize reward in the face of noisy, temporally evolving information. To test this theory, we trained an end-to-end deep recurrent neural network using reinforcement learning on a noisy perceptual discrimination task. Networks learned several key abilities of primate-like decision making including trading off speed for accuracy, and flexibly changing their mind in the face of new information. Internal dynamics of these networks suggest that these abilities were supported by similar decision mechanisms as those observed in primate neurophysiological studies. These results provide experimental support for key pressures that gave rise to the primate ability to make flexible decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。