arXiv:2608.04035cs.ARcs.AI2026-08中稿 · IEEE DFTS'26, 4 pa…

为视觉Transformer设计轻量级故障检测方法,提升可靠性

CheckOne: Lightweight Fault Detection and Mitigation for Vision Transformers

  • 采用对称校验机制,在不影响性能前提下检测故障
  • 相比传统方法,故障缓解能力提升26倍,平均性能高3.8倍
  • 适合部署在安全关键场景的视觉Transformer系统

视觉Transformer(ViTs)在安全关键应用中的广泛应用引发了硬件故障带来的可靠性问题。算法级容错(ABFT)方法作为轻量且对称的深度神经网络保护机制应运而生,但其在ViTs上应用面临显著计算开销挑战。本文全面评估了ViTs的可靠性,强调其各层需对称保护。为此提出CheckOne,一种新型、低成本的ViT故障检测与缓解方法,相比传统ABFT大幅降低计算开销。通过多组ViT模型的广泛实验验证,CheckOne在缓解关键故障方面能力提升达26倍,平均性能较ABFT高出3.8倍。

原文摘要 · Abstract (English)

The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns related to hardware faults. Algorithm-Based Fault Tolerance (ABFT) methods have emerged as lightweight and symmetric protection mechanisms for DNNs. However, they are particularly challenging for ViTs due to their significant computational requirements. This work comprehensively evaluates the reliability of ViTs, emphasizing the need for symmetric protection in their layers. Furthermore, we present CheckOne, a novel, cost-effective method for fault detection and mitigation in ViTs that significantly reduces the computational cost compared to conventional ABFT. Through extensive experiments with multiple ViTs, CheckOne mitigates critical faults by up to $26\times$ and achieves an average 3.8x higher performance than ABFT in ViTs.

视觉Transformer故障检测容错轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。