用噪声实现高效神经网络学习,适合低功耗硬件
Noise-based reward-modulated learning
- 利用神经活动噪声近似梯度,实现本地化学习
- 在多层网络中性能优于现有类似方法,收敛较慢但更稳定
- 适合事件驱动的类脑计算系统,尤其低功耗场景
追求能效与自适应的人工智能使类脑计算成为传统计算的有力替代。然而,在此类平台上实现学习需依赖以局部信息为主、有效信用分配的技术。本文提出基于噪声的奖励调制学习(NRL),一种将强化学习与基于梯度优化数学统一的突触可塑性规则,通过随机神经活动近似精确梯度,将生物与类脑平台固有的噪声转化为功能资源。受生物学习启发,该方法以奖励预测误差为目标生成渐进优势行为,并使用潜力痕迹实现回溯信用分配。在包含即时与延迟奖励的强化学习任务中,NRL表现接近反向传播优化基线,尽管收敛较慢;但在多层网络中显著优于最相近的方法——奖励调制赫布学习(RMHL),展现出更优性能与可扩展性。虽仅测试于简单架构,结果凸显了噪声驱动、脑启发学习在低功耗自适应系统中的潜力,特别适用于具有局部性约束的计算基底。NRL为下一代类脑人工智能的事件驱动特性提供了理论扎实的范式。
原文摘要 · Abstract (English)
The pursuit of energy-efficient and adaptive artificial intelligence (AI) has positioned neuromorphic computing as a promising alternative to conventional computing. However, achieving learning on these platforms requires techniques that prioritize local information while enabling effective credit assignment. Here, we propose noise-based reward-modulated learning (NRL), a novel synaptic plasticity rule that mathematically unifies reinforcement learning and gradient-based optimization with biologically-inspired local updates. NRL addresses the computational bottleneck of exact gradients by approximating them through stochastic neural activity, transforming the inherent noise of biological and neuromorphic substrates into a functional resource. Drawing inspiration from biological learning, our method uses reward prediction errors as its optimization target to generate increasingly advantageous behavior, and eligibility traces to facilitate retrospective credit assignment. Experimental validation on reinforcement tasks, featuring immediate and delayed rewards, shows that NRL achieves performance comparable to baselines optimized using backpropagation, although with slower convergence, while showing significantly superior performance and scalability in multi-layer networks compared to reward-modulated Hebbian learning (RMHL), the most prominent similar approach. While tested on simple architectures, the results highlight the potential of noise-driven, brain-inspired learning for low-power adaptive systems, particularly in computing substrates with locality constraints. NRL offers a theoretically grounded paradigm well-suited for the event-driven characteristics of next-generation neuromorphic AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。