arXiv:2601.19912cs.ARcs.AI2026-01

首次在指令级层面研究大模型推理对GPU软错误的脆弱性。

Analysis of LLM Vulnerability to GPU Soft Errors: An Instruction-Level Fault Injection Study

  • 通过指令级故障注入,分析大模型推理的可靠性特征。
  • 发现模型架构、参数规模和任务复杂度显著影响软错误敏感性。
  • 为大模型容错机制设计提供新依据,适合关注AI系统可靠性的研究者。

大型语言模型(LLMs)计算与内存需求极高,对高性能GPU提出严峻挑战。随着晶体管尺寸缩小和工作电压降低,GPU越来越容易受到软错误的影响。现有研究多聚焦通用应用或传统神经网络(如视觉分类与检测),对现代大规模LLMs的系统性分析仍显不足。鉴于LLMs的独特特性,其对软错误的鲁棒性可能与以往模型存在显著差异。为此,本文首次开展针对LLM推理的指令级故障注入研究,从模型架构、参数规模和任务复杂度等多个角度揭示其可靠性特征,为理解大模型可靠性提供新洞见,并指导更有效的容错机制设计。

原文摘要 · Abstract (English)

Large language models (LLMs) are highly compute- and memory-intensive, posing significant demands on high-performance GPUs. At the same time, advances in GPU technology driven by shrinking transistor sizes and lower operating voltages have made these devices increasingly susceptible to soft errors. While prior work has examined GPU reliability, most studies have focused on general-purpose applications or conventional neural networks mostly used for vision tasks such as classification and detection. In contrast, systematic analysis of modern large-scale LLMs remains limited, despite their rapid adoption in diverse application scenarios. Given the unique characteristics of LLMs, their resilience to soft errors may differ substantially from earlier models. To bridge this gap, we conduct the first instruction-level fault injection study of LLM inference. Our approach reveals reliability characteristics from multiple perspectives, highlighting the effects of model architecture, parameter scale, and task complexity. These findings provide new insights into LLM reliability and inform the design of more effective fault tolerance mechanisms.

大模型软错误GPU可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。