arXiv:2512.00427cs.ROphysics.optics2025-12

光子脉冲强化学习实现机器人连续控制,能效与延迟大幅优化。

Hardware-Software Collaborative Computing of Photonic Spiking Reinforcement Learning for Robotic Continuous Control

  • 光子芯片做矩阵运算,电子域完成脉冲激活,软硬协同计算。
  • 在HalfCheetah上获5831分奖励,收敛步数减少23.33%,动作偏差<2.2%。
  • 首次用可编程光子芯片做机器人控制,能效达1.39 TOPS/W,延迟仅120皮秒。

机器人连续控制任务因高维状态空间和实时交互需求,对计算架构的能效与延迟提出严苛要求。传统电子平台面临计算瓶颈,而光子计算与脉冲强化学习(RL)融合提供了新路径。本文提出基于光子脉冲强化学习的新架构,将双延迟深度确定性策略梯度(TD3)算法与脉冲神经网络(SNN)结合,采用光电混合计算范式:硅基光子马赫-曾德尔干涉仪(MZI)芯片执行线性矩阵运算,非线性脉冲激活在电子域完成。在Pendulum-v1与HalfCheetah-v2基准上的实验验证了软硬件协同推理能力,系统在HalfCheetah-v2上实现5831分控制奖励,收敛步数减少23.33%,动作偏差低于2.2%。该工作首次将可编程MZI光子计算芯片应用于机器人连续控制任务,达到1.39 TOPS/W的能效与120皮秒的超低计算延迟,展现了光子脉冲强化学习在自主与工业机器人实时决策中的巨大潜力。

原文摘要 · Abstract (English)

Robotic continuous control tasks impose stringent demands on the energy efficiency and latency of computing architectures due to their high-dimensional state spaces and real-time interaction requirements. Conventional electronic computing platforms face computational bottlenecks, whereas the fusion of photonic computing and spiking reinforcement learning (RL) offers a promising alternative. Here, we propose a novel computing architecture based on photonic spiking RL, which integrates the Twin Delayed Deep Deterministic policy gradient (TD3) algorithm with spiking neural network (SNN). The proposed architecture employs an optical-electronic hybrid computing paradigm wherein a silicon photonic Mach-Zehnder interferometer (MZI) chip executes linear matrix computations, while nonlinear spiking activations are performed in the electronic domain. Experimental validation on the Pendulum-v1 and HalfCheetah-v2 benchmarks demonstrates the system capability for software-hardware co-inference, achieving a control policy reward of 5831 on HalfCheetah-v2, a 23.33% reduction in convergence steps, and an action deviation below 2.2%. Notably, this work represents the first application of a programmable MZI photonic computing chip to robotic continuous control tasks, attaining an energy efficiency of 1.39 TOPS/W and an ultralow computational latency of 120 ps. Such performance underscores the promise of photonic spiking RL for real-time decision-making in autonomous and industrial robotic systems.

光子计算脉冲神经网络机器人控制能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。