针对芯片岛加速器内存访问瓶颈,提出自适应拓扑设计方法。
Taming the Tail: NoI Topology Synthesis for Mixed DL Workloads on Chiplet-Based Accelerators
- 基于干扰评分构建多目标优化模型,动态平衡性能与功耗。
- 生成拓扑使尾延迟降低至1.2倍,满足服务等级协议要求。
- 适合异构芯片岛系统中大规模模型推理场景的硬件设计。
异构芯片岛系统通过拆分CPU/GPU与新兴技术(如HBM/DRAM)实现可扩展性提升,但这种片上拆分引入了网络互连器(NoI)的延迟。我们观察到,在现代大模型推理中,参数与激活频繁在HBM/DRAM间往返移动,向互连器注入大量突发流量,导致尾延迟升高并违反基于k-ary n-cube架构的基准NoI拓扑的服务等级协议(SLAs)。为此,我们提出干扰评分(IS),量化竞争下的最坏情况延迟。将NoI拓扑设计建模为多目标优化问题,并开发PARL(分区感知强化学习器)拓扑生成器,以平衡吞吐量、延迟和功耗。所生成拓扑有效降低内存边界处的竞争,满足SLAs,将最坏情况延迟降至1.2倍,同时保持与链路丰富的网格拓扑相当的平均吞吐量。本工作重新定义了面向异构芯片岛加速器的工作负载感知式NoI设计范式。
原文摘要 · Abstract (English)
Heterogeneous chiplet-based systems improve scaling by disag-gregating CPUs/GPUs and emerging technologies (HBM/DRAM).However this on-package disaggregation introduces a latency inNetwork-on-Interposer(NoI). We observe that in modern large-modelinference, parameters and activations routinely move backand forth from HBM/DRAM, injecting large, bursty flows into theinterposer. These memory-driven transfers inflate tail latency andviolate Service Level Agreements (SLAs) across k-ary n-cube base-line NoI topologies. To address this gap we introduce an InterferenceScore (IS) that quantifies worst-case slowdown under contention.We then formulate NoI synthesis as a multi-objective optimization(MOO) problem. We develop PARL (Partition-Aware ReinforcementLearner), a topology generator that balances throughput, latency,and power. PARL-generated topologies reduce contention at the memory cut, meet SLAs, and cut worst-case slowdown to 1.2 times while maintaining competitive mean throughput relative to link-rich meshes. Overall, this reframes NoI design for heterogeneouschiplet accelerators with workload-aware objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。