对比多种生成模型在先进加速器上的性能,优化推理效率。
Performance Optimization and Comparative Analysis of Generative AI Models on Advanced Accelerators

- 提出混合精度后训练量化评估方法
- 在多类加速器上验证模型推理性能差异
- 适合硬件优化与模型部署研究者参考
生成式AI模型(如大语言模型和扩散模型)在众多任务中表现出色。然而,其部署面临内存占用高、推理延迟长、计算需求大及硬件成本高等挑战。这些难题在异构平台间尤为突出,因数值格式、内存带宽和软件栈差异与模型架构、工作负载特性相互作用,使性能评估复杂化。为此,本文系统研究了多种生成式AI模型在不同下游任务中的性能优化与对比分析,提出一种新型混合精度后训练量化评估方法,探索微调策略,并在现代高性能计算系统与先进加速器上评估性能表现。
原文摘要 · Abstract (English)
Generative AI models, such as Large Language Models (LLMs) and diffusion models, have demonstrated impressive performance across a wide range of tasks. Despite these advances, deployment remains challenging due to substantial memory requirements, extended inference latency, significant computational demands, and high hardware costs. These issues are further complicated when evaluating models across heterogeneous platforms, where differences in numerical formats, memory bandwidths, and software stacks interact with model architecture and workload characteristics in complex ways. To address these challenges, we present a systematic study focused on performance optimization and comparative analysis of several Generative AI models across diverse downstream tasks. This work introduces a novel mixed-precision post-training quantization evaluation, examines fine-tuning strategies, and assesses performance across modern high-performance computing (HPC) systems and advanced accelerators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。