arXiv:2504.06591cs.LGcs.SY2025-04中稿 · publication in IEE…被引 2

NAPER通过异构模型集成实现低延迟高可靠推理,解决资源受限场景下的故障容错难题。

NAPER: Fault Protection for Real-Time Resource-Constrained Deep Neural Networks

  • 采用异构模型冗余,多模型协同提升整体准确率
  • 比TMR方案快40%,故障下仍保持高精度(高出4.2%)
  • 适合实时性要求严苛的嵌入式智能系统

部署在资源受限系统的深度神经网络(DNN)在高精度应用中面临严格的时序要求,故障容错挑战显著。内存位翻转会严重降低DNN性能,而传统保护方法如三模冗余(TMR)常以牺牲准确性来保证可靠性,形成可靠性、准确性与及时性之间的三难困境。本文提出NAPER,一种基于集成学习的新型保护机制。不同于传统冗余方式,NAPER采用异构模型冗余,使多个不同模型联合预测优于任一单个模型。该方法结合高效的故障检测机制和实时调度器,能智能安排恢复操作,在不中断推理的前提下确保任务按时完成。评估表明:在正常及故障情况下,NAPER推理速度提升40%,准确性较基于TMR的策略高出4.2%,且故障恢复期间可保障持续运行。NAPER有效平衡了实时DNN应用中准确性、可靠性和及时性的多重需求。

原文摘要 · Abstract (English)

Fault tolerance in Deep Neural Networks (DNNs) deployed on resource-constrained systems presents unique challenges for high-accuracy applications with strict timing requirements. Memory bit-flips can severely degrade DNN accuracy, while traditional protection approaches like Triple Modular Redundancy (TMR) often sacrifice accuracy to maintain reliability, creating a three-way dilemma between reliability, accuracy, and timeliness. We introduce NAPER, a novel protection approach that addresses this challenge through ensemble learning. Unlike conventional redundancy methods, NAPER employs heterogeneous model redundancy, where diverse models collectively achieve higher accuracy than any individual model. This is complemented by an efficient fault detection mechanism and a real-time scheduler that prioritizes meeting deadlines by intelligently scheduling recovery operations without interrupting inference. Our evaluations demonstrate NAPER's superiority: 40% faster inference in both normal and fault conditions, maintained accuracy 4.2% higher than TMR-based strategies, and guaranteed uninterrupted operation even during fault recovery. NAPER effectively balances the competing demands of accuracy, reliability, and timeliness in real-time DNN applications

DNN容错实时系统异构冗余嵌入式推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。