arXiv:2606.15898cs.RO2026-06被引 1

用视觉语言模型指导脉冲神经网络,实现低功耗机器人视觉感知

VL2Spike: Spike-driven Distillation from VLMs for Low-Power Visual Perception in Embodied AI

论文配图:VL2Spike: Spike-driven Distillation from VLMs for Low-Power Visual Perception in Embodied AI
图 1 · 摘自论文原文
  • 通过时空脉冲对齐,将视觉语言模型知识迁移到脉冲网络
  • 在三个数据集上提升6.81%精度,能耗仅15.7%
  • 适合需要低功耗、强泛化的智能体视觉系统

脉冲神经网络(SNNs)是类脑的事件驱动模型,以稀疏脉冲计算,可在资源受限的具身智能体中实现高效视觉感知。近年来,具有脉冲自注意力的脉冲-变压器模型显著提升了纯SNN的学习能力。尽管SNN能效高,其性能仍受脉冲架构与优化挑战制约,标准梯度下降无法直接应用。视觉语言模型(VLMs)在多模态表征方面展现出丰富能力。为此,我们提出VL2Spike,一种新型脉冲知识蒸馏框架,将VLM的多模态知识与紧凑的Spikformer模型结合。该设计在保持能效优势的同时增强学习能力,为低功耗机器人感知提供可行路径。核心贡献包括:(1)空间-时间视觉脉冲(SVS)蒸馏,实现图像特征与脉冲标记间的共享流形对齐,以及膜电位与脉冲率的时序一致性预热;(2)脉冲原型引导的语言(SPL)蒸馏,对齐Spikformer的类别原型与提示式VLM文本嵌入。大量实验表明,VL2Spike在三个静态数据集上平均提升6.81%,能耗仅为15.7%。在机器人视觉位置识别(VPR)任务中也取得6.63%的增益,凸显其在具身智能低功耗感知中的潜力。

原文摘要 · Abstract (English)

Spiking neural networks (SNNs) are brain-inspired, event-driven models that compute with sparse spikes, which enables highly efficient visual perception in resource-constrained embodied AI models. The emergence of Spiking-Transformer models with spike self-attention has substantially improved the learning capacity of pure SNNs. Although SNNs are energy efficient, their performance is still limited by the spike-based architecture and optimization challenges, as standard gradient descent rules cannot be directly applied. Recently, vision-language models (VLMs) have shown rich multi-modal knowledge representation capabilities for visual perception. Thus, it is promising to leverage VLMs for better Spikformer training. To this end, we present VL2Spike, a novel spike-based knowledge distillation (KD) framework that bridges multi-modal knowledge from VLMs with compact Spikformer models. This design enhances the learning capacity of Spikformer models while preserving their energy-efficiency merits, thereby offering a practical pathway toward low-power robotic perception. Our VL2Spike brings two key technical contributions. To align with spiking dynamics, we first propose spatial-temporal visual spike (SVS) distillation, which achieves (1) shared manifold alignment between VLM image features and spike tokens, and (2) warm-started temporal consistency on membrane potentials and spike rates. We then design a novel spike prototype-guided linguistic (SPL) distillation strategy that aligns Spikformer's class prototypes and logits with promptable VLM text embeddings. Extensive experiments show that VL2Spike achieves 6.81% gain across three static datasets with only 15.7% energy consumption. It also exhibits strong generalization capacity on robotic visual place recognition (VPR) with a gain of 6.63%, highlighting its potential for low-power perception in embodied AI.

脉冲神经网络知识蒸馏低功耗具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。