PIM-AI通过内存计算架构,显著降低大模型推理的能耗与成本。
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
- 在内存芯片内集成计算单元,减少数据搬移瓶颈。
- 云场景下每秒查询成本降低6.94倍,移动端能耗降10至20倍。
- 适合追求能效与续航的云端及移动设备部署场景。
大型语言模型(LLMs)因其强大的语言理解与生成能力,在各类应用中日益重要。然而,其高计算与内存需求对传统硬件构成挑战。存算一体(PIM)通过将计算单元直接集成于内存芯片,可有效缓解数据传输瓶颈并提升能效。本文提出PIM-AI,一种针对DDR5/LPDDR5的新型PIM架构,专为LLM推理设计,无需修改内存控制器或物理层(PHY)。我们开发了仿真器,在多种场景下评估其性能,结果表明:在云端,基于不同模型,PIM-AI相较最先进GPU,3年总拥有成本(TCO)每查询降低最高达6.94倍;在移动端,相比最先进移动SoC,每令牌能耗降低10至20倍,查询速度提升25至45%,每查询能耗减少6.9至13.4倍,显著延长电池寿命,实现更多推理任务。这些成果凸显PIM-AI在提升模型部署效率、可扩展性与可持续性方面的巨大潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have become essential in a variety of applications due to their advanced language understanding and generation capabilities. However, their computational and memory requirements pose significant challenges to traditional hardware architectures. Processing-in-Memory (PIM), which integrates computational units directly into memory chips, offers several advantages for LLM inference, including reduced data transfer bottlenecks and improved power efficiency. This paper introduces PIM-AI, a novel DDR5/LPDDR5 PIM architecture designed for LLM inference without modifying the memory controller or DDR/LPDDR memory PHY. We have developed a simulator to evaluate the performance of PIM-AI in various scenarios and demonstrate its significant advantages over conventional architectures. In cloud-based scenarios, PIM-AI reduces the 3-year TCO per queries-per-second by up to 6.94x compared to state-of-the-art GPUs, depending on the LLM model used. In mobile scenarios, PIM-AI achieves a 10- to 20-fold reduction in energy per token compared to state-of-the-art mobile SoCs, resulting in 25 to 45~\% more queries per second and 6.9x to 13.4x less energy per query, extending battery life and enabling more inferences per charge. These results highlight PIM-AI's potential to revolutionize LLM deployments, making them more efficient, scalable, and sustainable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。