提出可重构容错架构,提升神经网络推理的可靠性与效率。
FORTALESA: Fault-Tolerant Reconfigurable Systolic Array for DNN Inference
- 动态切换三种执行模式,按层映射不同故障容忍策略。
- 对瞬态与永久故障有效保护,资源占用仅为静态冗余的1/6。
- 适合高可靠要求的自动驾驶、医疗等安全关键场景。
深度神经网络在任务和安全关键应用中的兴起使其可靠性成为核心问题。为满足高性能需求,专用硬件加速器被广泛采用。阵列结构因并行性与规则性,常用于神经网络加速器。本文提出一种运行时可重构的阵列架构,支持三种执行模式与四种实现方案。所有方案均从资源利用率、吞吐量及容错性能方面进行评估。通过异构映射不同网络层至不同执行模式,提升阵列上神经网络推理的可靠性。该方法基于故障传播分析的新型可靠性评估机制,用于探索最优的执行模式-层映射关系。所提架构能有效保护阵列处理单元中的寄存器与乘累加单元免受瞬态与永久故障影响。可重构特性使速度提升最高达3倍,取决于层的脆弱性;相比静态冗余,资源消耗减少6倍;相比先前针对瞬态故障的方案,资源减少2.5倍。
原文摘要 · Abstract (English)
The emergence of Deep Neural Networks (DNNs) in mission- and safety-critical applications brings their reliability to the front. High performance demands of DNNs require the use of specialized hardware accelerators. Systolic array architecture is widely used in DNN accelerators due to its parallelism and regular structure. This work presents a run-time reconfigurable systolic array architecture with three execution modes and four implementation options. All four implementations are evaluated in terms of resource utilization, throughput, and fault tolerance improvement. The proposed architecture is used for reliability enhancement of DNN inference on systolic array through heterogeneous mapping of different network layers to different execution modes. The approach is supported by a novel reliability assessment method based on fault propagation analysis. It is used for the exploration of the appropriate execution mode--layer mapping for DNN inference. The proposed architecture efficiently protects registers and MAC units of systolic array PEs from transient and permanent faults. The reconfigurability feature enables a speedup of up to $3\times$, depending on layer vulnerability. Furthermore, it requires $6\times$ fewer resources compared to static redundancy and $2.5\times$ fewer resources compared to the previously proposed solution for transient faults.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。