用强化学习优化大模型在3D CPU上的热管理与性能
Active Imitation Learning for Thermal- and Kernel-Aware LFM Inference on 3D S-NUCA Many-Cores
- 通过模仿学习自动学得最优调度策略
- 在保持温度安全前提下提升大模型推理性能
- 适合研究高性能多核系统与大模型部署的读者
大型基础模型(LFM)推理既耗内存又耗计算,传统依赖GPU,但受限于资源和成本,正转向高性能通用CPU,尤其是新兴的3D堆叠静态非均匀缓存架构(3D S-NUCA)系统。这类架构虽提供更高带宽和局部性,却因3D网络芯片(NoC)导致严重热问题和不均缓存延迟。由于LFM内核多样性和系统异构性,线程迁移与电压频率调节的最优管理极为复杂。现有热管理方法多依赖简化的分析模型,缺乏适应性。本文提出AILFM,一种基于主动模仿学习(AIL)的调度框架,从理想调度示范中学习近似最优的热感知调度策略,运行开销极低。AILFM同时考虑核心级性能差异和LFM内核特性,确保热安全的同时最大化性能。大量实验表明,AILFM优于当前最优基线,并在多种LFM工作负载上具有良好泛化能力。
原文摘要 · Abstract (English)
Large Foundation Model (LFM) inference is both memory- and compute-intensive, traditionally relying on GPUs. However, the limited availability and high cost have motivated the adoption of high-performance general-purpose CPUs, especially emerging 3D-stacked Static Non-Uniform Cache Architecture (3D S-NUCA) systems. These architectures offer enhanced bandwidth and locality but suffer from severe thermal challenges and uneven cache latencies due to 3D Networks-on-Chip (NoC). Optimal management of thread migration and V/f scaling is non-trivial due to LFM kernel diversity and system heterogeneity. Existing thermal management approaches often rely on oversimplified analytical models and lack adaptability. We propose AILFM, an Active Imitation Learning (AIL)-based scheduling framework that learns near-optimal thermal-aware scheduling policies from Oracle demonstrations with minimal run-time overhead. AILFM accounts for both core-level performance heterogeneity and kernel-specific behavior in LFMs to maintain thermal safety while maximizing performance. Extensive experiments show that AILFM outperforms state-of-the-art baselines and generalizes well across diverse LFM workloads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。