测试NVIDIA Hopper GPU启用可信执行环境对大模型推理的性能影响
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
- 在Hopper GPU上启用可信执行环境,评估其对大模型推理的开销
- 主要瓶颈是CPU-GPU通过PCIe的数据传输,整体开销普遍低于7%
- 大模型和长序列任务几乎无额外开销,适合高安全需求场景
本报告评估了在NVIDIA Hopper GPU上启用可信执行环境(TEE)对大语言模型(LLM)推理任务的性能影响。我们在不同LLM和不同令牌长度下进行了基准测试,重点关注通过PCIe进行的CPU-GPU数据传输带来的瓶颈。结果显示,尽管GPU内部计算开销极小,但整体性能下降主要源于数据传输延迟。对于大多数典型LLM查询,开销低于7%;大模型和长序列任务的开销几乎可忽略不计。
原文摘要 · Abstract (English)
This report evaluates the performance impact of enabling Trusted Execution Environments (TEE) on NVIDIA Hopper GPUs for large language model (LLM) inference tasks. We benchmark the overhead introduced by TEE mode across various LLMs and token lengths, with a particular focus on the bottleneck caused by CPU-GPU data transfers via PCIe. Our results indicate that while there is minimal computational overhead within the GPU, the overall performance penalty is primarily attributable to data transfer. For the majority of typical LLM queries, the overhead remains below 7%, with larger models and longer sequences experiencing nearly zero overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。