arXiv:2606.27807cs.RO2026-06中稿 · ICML被引 1

用脉冲神经网络实现低功耗智能体导航,兼顾效率与性能。

SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks

论文配图:SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks
图 1 · 摘自论文原文
  • 采用事件驱动的脉冲神经网络替代传统密集模型,降低视觉表征能耗。
  • 在导航与机器人控制任务中,能耗和计算成本显著下降,性能仍具竞争力。
  • 适合对功耗敏感的实时嵌入式智能系统,如移动机器人、可穿戴设备。

视觉-语言-动作(VLA)模型已成为具身智能的主流范式,但多数基于大规模Transformer架构,导致推理延迟高、能耗大,限制了其在低功耗、实时场景中的部署。本文提出SpikeVLA,一种面向具身导航的脉冲神经网络VLA架构,包含三个核心组件:(i) 脉冲视觉编码器Spike-V,以事件驱动的脉冲层替代密集连续层,降低视觉表征学习能耗;(ii) 多模态脉冲大语言模型Spike-L,通过脉冲动力学重构跨模态推理,并引入令牌级事件驱动稀疏性以进一步降低计算开销;(iii) 脉冲动作策略网络Spike-A,采用拉普拉斯核群体编码与多层全连接脉冲神经网络,将脉冲活动解码为稳定可靠的连续控制信号,在低功耗约束下实现高效推理。在导航与机器人控制任务上的实验表明,SpikeVLA显著降低了能耗与计算成本,同时保持了竞争性性能,展现出在低功耗、实时具身智能中的应用潜力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have become a dominant paradigm for embodied intelligence. However, most existing approaches are built on large-scale transformers, resulting in substantial inference latency and energy consumption that limit their practical deployment in low-power, real-time scenarios. We propose SpikeVLA, a spiking VLA architecture for embodied navigation with energy-efficient inference, consisting of three key components. (i) A spiking vision encoder, Spike-V, that replaces dense continuous layers with event-driven spiking layers to reduce the energy consumption of visual representation learning. (ii) A multi-modal spiking large language model, Spike-L, that reformulates cross-modal reasoning with spiking dynamics and token-level event-driven sparsity to further lower computational cost. (iii) A spiking action policy network, Spike-A employs Laplacian-kernel population coding with a multi-layer fully connected SNN, and decodes spiking activities into stable and robust continuous control with energy-efficient inference under low-power constraints. Experiments on navigation and robotic control tasks show that SpikeVLA significantly reduces energy consumption and computational cost while maintaining competitive performance, highlighting its potential for low-power, real-time embodied intelligence.

脉冲神经网络具身智能低功耗多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。