用强化学习自动找到大模型加速器最致命的故障场景,提速超2倍、测试量降99%。
RIFT: A Scalable Methodology for LLM Accelerator Fault Assessment using Reinforcement Learning
- 将故障探测转为序列决策问题,结合敏感性分析与强化学习搜索最优故障
- 在千亿参数大模型上实现2.2倍加速,测试数据量减少超99%,覆盖更优
- 生成可直接用于商业验证流程的测试文件,适合芯片设计与可靠性团队
现代人工智能加速器规模巨大,传统故障评估方法面临计算成本过高和关键故障模式覆盖不足的问题。本文提出RIFT(强化学习引导的智能故障定位),一种可扩展框架,能自动发现最小且高影响的故障场景,实现高效的设计阶段故障评估。RIFT将复杂故障搜索转化为序列决策问题,结合混合敏感性分析进行搜索空间剪枝,并利用强化学习智能生成最小、高影响的测试集。在使用NVIDIA A100 GPU的千亿参数大语言模型工作负载上评估,RIFT相比进化方法实现2.2倍的故障评估速度提升,测试向量数量减少超过99%,同时达到更优的故障覆盖率。该框架还提供可行动数据,支持智能硬件保护策略:基于RIFT指导的定向纠错码比均匀三模冗余方案在单位面积上的覆盖率提升12.8倍,显著提高成本效益。RIFT自动生成符合UVM标准的验证产物,确保结果可直接用于商业RTL验证流程。
原文摘要 · Abstract (English)
The massive scale of modern AI accelerators presents critical challenges to traditional fault assessment methodologies, which face prohibitive computational costs and provide poor coverage of critical failure modes. This paper introduces RIFT (Reinforcement Learning-guided Intelligent Fault Targeting), a scalable framework that automates the discovery of minimal, high-impact fault scenarios for efficient design-time fault assessment. RIFT transforms the complex search for worst-case faults into a sequential decision-making problem, combining hybrid sensitivity analysis for search space pruning with reinforcement learning to intelligently generate minimal, high-impact test suites. Evaluated on billion-parameter Large Language Model (LLM) workloads using NVIDIA A100 GPUs, RIFT achieves a \textbf{2.2$\times$} fault assessment speedup over evolutionary methods and reduces the required test vector volume by over \textbf{99\%} compared to random fault injection, all while achieving \textbf{superior fault coverage}. The proposed framework also provides actionable data to enable intelligent hardware protection strategies, demonstrating that RIFT-guided selective error correction code provides a \textbf{12.8$\times$} improvement in \textbf{cost-effectiveness} (coverage per unit area) compared to uniform triple modular redundancy protection. RIFT automatically generates UVM-compliant verification artifacts, ensuring its findings are directly actionable and integrable into commercial RTL verification workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。