arXiv:2507.13736cs.LGcs.AR2025-07被引 3

将大模型直接部署到神经形态芯片上运行,实现端到端推理

An End-to-End DNN Inference Framework for the SpiNNaker2 Neuromorphic MPSoC

  • 基于OctopuScheduler扩展多层调度框架,打通从PyTorch到芯片的全流程
  • 支持高达Transformer规模的复杂DNN在单颗SpiNNaker2芯片上运行
  • 适合边缘计算场景下低功耗、高能效的神经网络推理任务

本文提出一种多层深度神经网络调度框架,作为OctopuScheduler的扩展,实现了从PyTorch模型到单个SpiNNaker2芯片上推理的端到端流程。结合包含量化和降低步骤的前端处理,该框架使大型复杂DNN(最大可达Transformer规模)能够在神经形态平台SpiNNaker2上进行边缘执行。

原文摘要 · Abstract (English)

This work presents a multi-layer DNN scheduling framework as an extension of OctopuScheduler, providing an end-to-end flow from PyTorch models to inference on a single SpiNNaker2 chip. Together with a front-end comprised of quantization and lowering steps, the proposed framework enables the edge-based execution of large and complex DNNs up to transformer scale using the neuromorphic platform SpiNNaker2.

神经形态计算边缘推理DNN调度SpiNNaker2

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。