arXiv:2605.24016cs.ARcs.AI2026-05中稿 · paper

首个专用于局部耦合相位扩散的低功耗硬件加速器,显著提升边缘计算效率。

SA-Kura: An Energy-Efficient Systolic Array Accelerator for Locally-Coupled Kuramoto Drift in Diffusion Sampling

论文配图:SA-Kura: An Energy-Efficient Systolic Array Accelerator for Locally-Coupled Kuramoto Drift in Diffusion Sampling
图 1 · 摘自论文原文
  • 将非线性相位耦合重构成可流水线执行的规则计算模式
  • 相比处理器软件运行,延迟降低193倍,能效提升69.4倍
  • 适合边缘设备上高效实现复杂扩散采样,尤其适用于资源受限场景

扩散采样在边缘部署中仍成本高昂,现有加速器几乎只关注得分网络,因标准漂移仅为简单线性缩放。局部耦合的柯朗多相位扩散以邻域相位交互替代该线性漂移,提升了采样效率,但引入新硬件瓶颈:每一步反向传播均需计算中心依赖的非线性5×5模板。该核函数难以适配传统CNN加速器与矩阵导向引擎。本文提出SA-Kura,据我们所知首个针对局部耦合柯朗多漂移设计的数字脉动阵列加速器。通过将成对正弦耦合重构为不依赖中心相位的邻域累加,再结合单次中心依赖乘减操作,消除阵列内超越函数单元,实现寄存器级复用与规则脉动执行。SA-Kura采用可综合RTL实现,集成于轻量级RISC-V SoC,FPGA原型验证,并经45 nm CMOS综合与功耗分析。仅针对漂移核,相比同平台处理器核心的软件实现,延迟降低193倍,能效提升69.4倍;相较独立的Jetson Orin Nano CUDA实现,速度提升6.57倍,每像素能耗降低约46.0倍。

原文摘要 · Abstract (English)

Diffusion inference remains costly for edge deployment, yet existing accelerators focus almost exclusively on score networks because standard drift is merely a trivial linear scaling. Kuramoto orientation diffusion replaces this trivial drift with locally coupled phase interactions, improving sampling efficiency but introducing a new hardware bottleneck: a center-dependent nonlinear 5 x 5 stencil evaluated at every reverse step. This kernel maps poorly to conventional CNN accelerators and matrix-oriented engines. We present SA-Kura, to our knowledge the first digital systolic-array accelerator dedicated to locally coupled Kuramoto drift. By reformulating pair-wise sinusoidal coupling into neighbor accumulation independent of the center phase followed by a single center-dependent multiply-subtract combination, SA-Kura eliminates in-PE transcendental units and enables regular systolic execution with register-level reuse. SA-Kura was implemented in synthesizable RTL, integrated into a lightweight RISC-V-based SoC, prototyped on FPGA, and evaluated through 45 nm CMOS synthesis and power analysis. For the drift kernel only, compared with software execution of the same kernel on the processor core in the same SoC platform, SA-Kura reduces latency and energy by 193x and 69.4x, respectively. Compared with a standalone Jetson Orin Nano CUDA implementation of the same kernel, it is 6.57x faster and achieves approximately 46.0x lower energy per pixel.

边缘计算硬件加速扩散模型脉动阵列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。