让手机电脑跑更大生成模型,性能提升十倍。
Scaling On-Device GPU Inference for Large Generative Models
- 设计新框架ML Drift,优化GPU推理能力。
- 可在设备端运行参数量高10到100倍的生成模型。
- 适配手机和电脑多平台,适合隐私敏感场景使用。
随着生成式AI的发展,大型机器学习模型已革新图像处理、音频合成和语音识别等领域。尽管服务器部署仍具性能优势,但出于隐私与效率考量,设备端推理需求持续存在。鉴于GPU是目前覆盖面最广的设备级AI加速器,我们提出ML Drift——一个优化的框架,扩展了当前先进GPU加速推理引擎的能力。该框架使设备端可执行参数量比现有模型高10至100倍的生成式AI任务。ML Drift解决了跨GPU API开发的复杂工程挑战,确保在移动及桌面/笔记本平台上的广泛兼容性,从而实现更复杂模型在资源受限设备上的部署。我们的GPU加速推理引擎相较现有开源方案,性能提升一个数量级。
原文摘要 · Abstract (English)
Driven by the advancements in generative AI, large machine learning models have revolutionized domains such as image processing, audio synthesis, and speech recognition. While server-based deployments remain the locus of peak performance, the imperative for on-device inference, necessitated by privacy and efficiency considerations, persists. Recognizing GPUs as the on-device ML accelerator with the widest reach, we present ML Drift--an optimized framework that extends the capabilities of state-of-the-art GPU-accelerated inference engines. ML Drift enables on-device execution of generative AI workloads which contain 10 to 100x more parameters than existing on-device generative AI models. ML Drift addresses intricate engineering challenges associated with cross-GPU API development, and ensures broad compatibility across mobile and desktop/laptop platforms, thereby facilitating the deployment of significantly more complex models on resource-constrained devices. Our GPU-accelerated ML/AI inference engine achieves an order-of-magnitude performance improvement relative to existing open-source GPU inference engines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。