评测软件化GPU虚拟化系统,助你选对云上算力方案
GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
- 构建56项指标的综合评测框架,覆盖性能、隔离、调度等10类场景
- 对比HAMI-core、BUD-FCSP与模拟MIG,揭示真实部署性能差异
- 专为多租户云环境设计,适合云计算和AI推理资源规划者
GPU加速计算在人工智能和大语言模型(LLM)推理中的普及,催生了云和容器环境中高效共享GPU资源的迫切需求。尽管NVIDIA的多实例GPU(MIG)技术提供硬件级隔离,但仅限高端数据中心GPU。HAMI-core和BUD-FCSP等软件虚拟化方案可支持更广泛GPU系列,但缺乏标准化评估方法。本文提出GPU-Virt-Bench,一个涵盖56项性能指标、分为10个类别的综合性基准测试框架,评估虚拟化系统的开销、隔离质量、LLM性能、内存带宽、缓存行为、PCIe吞吐、多GPU通信、调度效率、内存碎片及错误恢复能力。该框架支持软件虚拟化方案与理想MIG行为的系统性对比,为多租户环境下GPU资源部署提供决策依据。我们通过评估HAMI-core、BUD-FCSP及模拟MIG基线验证了其有效性,揭示了影响生产部署的关键性能特征。
原文摘要 · Abstract (English)
The proliferation of GPU-accelerated workloads, particularly in artificial intelligence and large language model (LLM) inference, has created unprecedented demand for efficient GPU resource sharing in cloud and container environments. While NVIDIA's Multi-Instance GPU (MIG) technology provides hardware-level isolation, its availability is limited to high-end datacenter GPUs. Software-based virtualization solutions such as HAMi-core and BUD-FCSP offer alternatives for broader GPU families but lack standardized evaluation methodologies. We present GPU-Virt-Bench, a comprehensive benchmarking framework that evaluates GPU virtualization systems across 56 performance metrics organized into 10 categories. Our framework measures overhead, isolation quality, LLM-specific performance, memory bandwidth, cache behavior, PCIe throughput, multi-GPU communication, scheduling efficiency, memory fragmentation, and error recovery. GPU-Virt-Bench enables systematic comparison between software virtualization approaches and ideal MIG behavior, providing actionable insights for practitioners deploying GPU resources in multi-tenant environments. We demonstrate the framework's utility through evaluation of HAMi-core, BUD-FCSP, and simulated MIG baselines, revealing performance characteristics critical for production deployment decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。