小误差引发大故障,提出三招提升存算一体芯片可靠性
When Small Variations Become Big Failures: Reliability Challenges in Compute-in-Memory Neural Accelerators
- 通过选择性写入验证,精准定位并修复易错区域
- 微小器件波动可致推理精度骤降,甚至灾难性失效
- 训练时加入真实噪声模拟,无需额外硬件即可增强鲁棒性
存算一体(CiM)架构通过缓解冯·诺依曼瓶颈,显著提升深度神经网络加速的能效与吞吐量。然而,其依赖新兴非易失性存储器件,引入了写入波动、电导漂移和随机噪声等器件级非理想特性,严重挑战可靠性、可预测性与安全性,尤其在安全关键场景中。本研究揭示:即使微小的器件变化,也会在安全关键推理任务中引发不成比例的精度下降乃至灾难性失败,暴露出平均性能评估与最坏情况行为之间的巨大差距。基于此,提出SWIM机制,仅在最关键位置实施写入验证,大幅提升可靠性同时保持高效性。此外,提出一种以学习为中心的解决方案:在训练中注入右截断高斯噪声,使训练假设贴近硬件实际波动,实现无额外硬件开销的鲁棒部署。这些工作强调跨层协同设计的重要性,为在新兴内存技术上实现可靠、高效的神经推理提供了系统性路径,推动其在安全与可靠性要求高的系统中的应用。
原文摘要 · Abstract (English)
Compute-in-memory (CiM) architectures promise significant improvements in energy efficiency and throughput for deep neural network acceleration by alleviating the von Neumann bottleneck. However, their reliance on emerging non-volatile memory devices introduces device-level non-idealities-such as write variability, conductance drift, and stochastic noise-that fundamentally challenge reliability, predictability, and safety, especially in safety-critical applications. This talk examines the reliability limits of CiM-based neural accelerators and presents a series of techniques that bridge device physics, architecture, and learning algorithms to address these challenges. We first demonstrate that even small device variations can lead to disproportionately large accuracy degradation and catastrophic failures in safety-critical inference workloads, revealing a critical gap between average-case evaluations and worst-case behavior. Building on this insight, we introduce SWIM, a selective write-verify mechanism that strategically applies verification only where it is most impactful, significantly improving reliability while maintaining CiM's efficiency advantages. Finally, we explore a learning-centric solution that improves realistic worst-case performance by training neural networks with right-censored Gaussian noise, aligning training assumptions with hardware-induced variability and enabling robust deployment without excessive hardware overhead. Together, these works highlight the necessity of cross-layer co-design for CiM accelerators and provide a principled path toward dependable, efficient neural inference on emerging memory technologies-paving the way for their adoption in safety- and reliability-critical systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。