arXiv:2511.06770cs.AReess.IV2025-11被引 2

为事件驱动视觉推理设计低功耗脉冲变压器加速器,兼顾能效与实时性。

ASTER: Attention-based Spiking Transformer Engine for Event-driven Reasoning

  • 采用存内计算架构,结合模拟数字混合设计优化稀疏输入
  • 在图像与事件数据集上实现比GPU低467倍的能耗
  • 适合边缘设备上实时、低功耗的智能视觉应用

将脉冲神经网络(SNN)与基于Transformer的架构融合,为边缘设备上的生物启发式低功耗事件驱动视觉推理开辟新路径。然而,脉冲计算的高时间分辨率和二值特性与传统数字硬件(如CPU/GPU)存在架构不匹配问题。现有类脑与存算一体(PIM)加速器难以应对此类模型中普遍存在的高稀疏性和复杂运算。为此,我们提出一种面向脉冲Transformer的内存中心型硬件加速器,专为实时事件驱动框架(如静态与事件输入帧分类)部署而优化。设计采用混合模拟-数字PIM架构,引入输入稀疏性优化,并设计定制数据流,以最小化内存访问开销并最大化时空稀疏条件下的数据复用,实现计算与内存高效的端到端脉冲Transformer执行。进一步提出推理时软件优化策略,包括层跳过与时间步缩减,利用贝叶斯优化与代理建模,在严格计算预算下对算法-微架构联合空间进行鲁棒高效协同探索。在图像(ImageNet)与事件数据集(CIFAR-10 DVS、DVSGesture)上评估,该加速器相较边缘GPU(Jetson Orin Nano)实现最高约467倍、较先前脉冲变压器PIM加速器提升1.86倍的能效,同时在ImageNet上保持竞争力的任务准确率。本工作推动了一类新型普适边缘人工智能的发展,基于脉冲变压器加速实现在极端边缘的低功耗、实时视觉处理。

原文摘要 · Abstract (English)

The integration of spiking neural networks (SNNs) with transformer-based architectures has opened new opportunities for bio-inspired low-power, event-driven visual reasoning on edge devices. However, the high temporal resolution and binary nature of spike-driven computation introduce architectural mismatches with conventional digital hardware (CPU/GPU). Prior neuromorphic and Processing-in-Memory (PIM) accelerators struggle with high sparsity and complex operations prevalent in such models. To address these challenges, we propose a memory-centric hardware accelerator tailored for spiking transformers, optimized for deployment in real-time event-driven frameworks such as classification with both static and event-based input frames. Our design leverages a hybrid analog-digital PIM architecture with input sparsity optimizations, and a custom-designed dataflow to minimize memory access overhead and maximize data reuse under spatiotemporal sparsity, for compute and memory-efficient end-to-end execution of spiking transformers. We subsequently propose inference-time software optimizations for layer skipping, and timestep reduction, leveraging Bayesian Optimization with surrogate modeling to perform robust, efficient co-exploration of the joint algorithmic-microarchitectural design spaces under tight computational budgets. Evaluated on both image(ImageNet) and event-based (CIFAR-10 DVS, DVSGesture) classification, the accelerator achieves up to ~467x and ~1.86x energy reduction compared to edge GPU (Jetson Orin Nano) and previous PIM accelerators for spiking transformers, while maintaining competitive task accuracy on ImageNet dataset. This work enables a new class of intelligent ubiquitous edge AI, built using spiking transformer acceleration for low-power, real-time visual processing at the extreme edge.

脉冲神经网络存算一体边缘计算Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。