arXiv:2409.16633cs.ARcs.DC2024-09被引 8

通过硬件开关近数据处理,显著加速大规模推荐系统推理。

PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System Inferences

  • 在CXL交换机下游端口实现近数据计算,减少数据搬移延迟。
  • 相比主流方案Pond,推理延迟降低3.89倍;比BEACON快2.03倍。
  • 适合数据中心高性能推荐系统优化,尤其关注带宽与内存扩展性。

深度学习推荐模型(DLRMs)已成为现代数据中心中AI推理的主要负载,其性能受嵌入表向量规模大及并发访问的影响显著。为突破现有方案瓶颈,亟需结合新兴互连技术如CXL的新优化方法。本研究深入分析了运行于支持CXL系统的工业级DLRM工作负载,识别出当前CXL系统中的主要瓶颈。为此,提出PIFS-Rec,一种基于过程内硬件开关(PIFS)的方案,在互连开关的下游端口实现近数据处理,以加速推理并提升内存与带宽可扩展性。实验表明,PIFS-Rec的延迟较行业标准的Pond系统低3.89倍,且相较先进方案BEACON提升2.03倍。

原文摘要 · Abstract (English)

Deep Learning Recommendation Models (DLRMs) have become increasingly popular and prevalent in today's datacenters, consuming most of the AI inference cycles. The performance of DLRMs is heavily influenced by available bandwidth due to their large vector sizes in embedding tables and concurrent accesses. To achieve substantial improvements over existing solutions, novel approaches towards DLRM optimization are needed, especially, in the context of emerging interconnect technologies like CXL. This study delves into exploring CXL-enabled systems, implementing a process-in-fabric-switch (PIFS) solution to accelerate DLRMs while optimizing their memory and bandwidth scalability. We present an in-depth characterization of industry-scale DLRM workloads running on CXL-ready systems, identifying the predominant bottlenecks in existing CXL systems. We, therefore, propose PIFS-Rec, a PIFS-based scheme that implements near-data processing through downstream ports of the fabric switch. PIFS-Rec achieves a latency that is 3.89x lower than Pond, an industry-standard CXL-based system, and also outperforms BEACON, a state-of-the-art scheme, by 2.03x.

推荐系统CXL近数据处理硬件优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。