arXiv:2503.11901cs.DCcs.AI2025-03被引 17

对比H100与A100 GPU在真实系统中的可靠性,发现内存更脆弱但硬件整体更稳。

Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs

  • 基于1170万小时数据,分析两种GPU在超大规模系统中的错误表现
  • H100内存平均无故障时间比A100低3.2倍,恢复机制不足
  • 虽内存易出错,但关键硬件组件可靠性显著优于A100,适合大规模部署

本研究基于包含1,056块A100和H100 GPU、峰值算力超1,300 petaflops的大型AI系统Delta,分析了长达2.5年的运行数据(共1170万GPU小时),揭示多项关键发现:(i) H100 GPU的内存错误平均无故障时间(MTBE)为A100的3.2倍,表明其内存可靠性更差;(ii) H100内存错误恢复机制无法应对更高内存容量带来的压力;(iii) 在关键硬件组件方面,H100整体硬件韧性显著优于A100;(iv) A100与H100的GPU错误常导致任务失败,因应用层缺乏稳健恢复机制;(v) 预测大规模系统中需预留5%以上的冗余节点以应对故障。该研究为下一代异构计算系统的容错设计提供实证依据。

原文摘要 · Abstract (English)

This study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years of operational data (11.7 million GPU hours) on GPU errors. Our major findings include: (i) H100 GPU memory resilience is worse than A100 GPU memory, with 3.2x lower per-GPU MTBE for memory errors, (ii) The GPU memory error-recovery mechanisms on H100 GPUs are insufficient to handle the increased memory capacity, (iii) H100 GPUs demonstrate significantly improved GPU hardware resilience over A100 GPUs with respect to critical hardware components, (iv) GPU errors on both A100 and H100 GPUs frequently result in job failures due to the lack of robust recovery mechanisms at the application level, and (v) We project the impact of GPU node availability on larger-scales and find that significant overprovisioning of 5% is necessary to handle GPU failures.

GPU可靠性大模型运维硬件容错

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。