用元学习发现可实现延迟奖励学习的局部突触规则。
Meta-learning three-factor plasticity rules for structured credit assignment with sparse feedback
- 通过元学习优化局部突触可塑性参数,结合任务执行与外部优化循环。
- 仅用局部信息和延迟奖励,实现长时间尺度的信用分配。
- 为生物合理神经网络学习机制提供新思路,适合类脑计算研究者。
生物神经网络能从稀疏、延迟的反馈中学习复杂行为,依赖局部突触可塑性,但其结构化信用分配机制仍不明确。相比之下,人工递归网络解决类似任务通常依赖非生物学的全局学习规则或手工设计的局部更新。支持延迟强化学习的局部可塑性规则空间尚未被充分探索。本文提出一种元学习框架,用于在稀疏反馈下发现递归网络中支持结构化信用分配的局部学习规则。方法在任务执行中采用类新海布式局部更新,并通过外层循环利用梯度传播优化可塑性参数。所得三因素学习规则仅依赖局部信息和延迟奖励,即可实现长时序信用分配,为递归电路中的生物合理学习机制提供了新见解。
原文摘要 · Abstract (English)
Biological neural networks learn complex behaviors from sparse, delayed feedback using local synaptic plasticity, yet the mechanisms enabling structured credit assignment remain elusive. In contrast, artificial recurrent networks solving similar tasks typically rely on biologically implausible global learning rules or hand-crafted local updates. The space of local plasticity rules capable of supporting learning from delayed reinforcement remains largely unexplored. Here, we present a meta-learning framework that discovers local learning rules for structured credit assignment in recurrent networks trained with sparse feedback. Our approach interleaves local neo-Hebbian-like updates during task execution with an outer loop that optimizes plasticity parameters via \textbf{tangent-propagation through learning}. The resulting three-factor learning rules enable long-timescale credit assignment using only local information and delayed rewards, offering new insights into biologically grounded mechanisms for learning in recurrent circuits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。