联合优化硬件与工作负载,提升通用存内计算芯片能效。
Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
- 通过联合优化确定最佳硬件参数,兼顾多种工作负载。
- 相比单一最大负载优化,能效-延迟-面积综合得分提升20%至69%。
- 适合需要多任务兼容的存内计算芯片设计者参考。
设计能高效支持多种工作负载的通用存内计算(IMC)硬件需要大量设计空间探索,手动完成不可行。单独为每个工作负载或仅针对最大工作负载优化硬件,往往无法获得最优通用解。为此,我们提出一种联合硬件-工作负载优化框架,可识别出最优的IMC芯片架构参数,实现更高效且负载灵活的硬件设计。实验表明,与仅针对单一最大工作负载优化的方案相比,该方法在VGG16、ResNet18、AlexNet和MobileNetV3上的能量-延迟-面积综合得分分别提升36%、36%、20%和69%。此外,我们量化了通用IMC硬件相较于特定工作负载设计的性能权衡与损失。
原文摘要 · Abstract (English)
Designing generalized in-memory computing (IMC) hardware that efficiently supports a variety of workloads requires extensive design space exploration, which is infeasible to perform manually. Optimizing hardware individually for each workload or solely for the largest workload often fails to yield the most efficient generalized solutions. To address this, we propose a joint hardware-workload optimization framework that identifies optimised IMC chip architecture parameters, enabling more efficient, workload-flexible hardware. We show that joint optimization achieves 36%, 36%, 20%, and 69% better energy-latency-area scores for VGG16, ResNet18, AlexNet, and MobileNetV3, respectively, compared to the separate architecture parameters search optimizing for a single largest workload. Additionally, we quantify the performance trade-offs and losses of the resulting generalized IMC hardware compared to workload-specific IMC designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。