将随机数生成器嵌入内存,实现低功耗神经网络不确定性计算
A 65 nm Bayesian Neural Network Accelerator with 360 fJ/Sample In-Word GRNG for AI Uncertainty Estimation
- 把高精度随机数生成器直接做进SRAM存储单元,减少延迟和能耗
- 每样本仅耗电360飞焦,实现5.12吉采样率的随机数输出
- 适合自动驾驶、医疗诊断等需实时判断不确定性的边缘AI场景
不确定性估计是自动驾驶、医疗诊断等安全关键型AI应用不可或缺的能力。贝叶斯神经网络(BNNs)利用贝叶斯统计可同时提供分类结果与不确定性评估,但其因频繁的随机数生成和重复采样带来高昂计算开销。此外,由于每次随机数生成后需频繁写入内存,BNN难以直接适配存内计算架构。为此,本文提出一款ASIC芯片,将360 fJ/样本的高精度高斯随机数生成器(GRNG)直接集成于SRAM存储单元中。该设计显著降低随机数生成开销,并支持完全并行的存内计算,实现高效推理。原型芯片面积仅0.45 mm²,达成5.12 GSa/s的随机数生成吞吐量和102 GOp/s的神经网络运算吞吐量,使边缘设备具备实时不确定性估计能力。
原文摘要 · Abstract (English)
Uncertainty estimation is an indispensable capability for AI-enabled, safety-critical applications, e.g. autonomous vehicles or medical diagnosis. Bayesian neural networks (BNNs) use Bayesian statistics to provide both classification predictions and uncertainty estimation, but they suffer from high computational overhead associated with random number generation and repeated sample iterations. Furthermore, BNNs are not immediately amenable to acceleration through compute-in-memory architectures due to the frequent memory writes necessary after each RNG operation. To address these challenges, we present an ASIC that integrates 360 fJ/Sample Gaussian RNG directly into the SRAM memory words. This integration reduces RNG overhead and enables fully-parallel compute-in-memory operations for BNNs. The prototype chip achieves 5.12 GSa/s RNG throughput and 102 GOp/s neural network throughput while occupying 0.45 mm2, bringing AI uncertainty estimation to edge computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。