arXiv:2511.04321cs.ARcs.AI2025-11中稿 · ISCA 2025被引 2

通过软硬件协同设计,显著缓解高性能存算芯片的电压降问题。

AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM

  • 利用存算架构特性建立工作负载与电压降的直接关联。
  • 实现最高69.2%电压降抑制,能效提升2.29倍,速度加快1.152倍。
  • 适合关注高密度存算芯片能效与可靠性的研究人员和工程师。

SRAM存算一体(PIM)已成为高性能存算最具前景的实现方式,具备更高的计算密度、能效与精度。但为追求更高性能,电路设计更复杂、运行频率更高,加剧了电压降(IR-drop)问题。严重电压降会显著降低芯片性能,威胁可靠性。传统电路级缓解方法如后端优化资源开销大,常牺牲功耗、性能与面积(PPA)。为此,本文提出AIM,一种面向高性能存算的全面软硬件协同架构级电压降缓解方案。首先,基于存算的位串行与就地数据流特性,提出Rtog与HR,建立存算负载与电压降的直接关联。在此基础上,提出LHR与WDS,通过软件优化实现架构级电压降缓解,保障计算精度。随后,设计IR-Booster动态调节机制,融合软件层HR信息与硬件级电压降监控,自适应调整存算宏的电压-频率配对,提升能效与性能。最后,提出HR感知的任务映射方法,打通软硬件设计,实现最优效果。7nm 256-TOPS存算芯片后布局仿真结果表明,AIM可实现最高69.2%电压降缓解,带来2.29倍能效提升与1.152倍加速。

原文摘要 · Abstract (English)

SRAM Processing-in-Memory (PIM) has emerged as the most promising implementation for high-performance PIM, delivering superior computing density, energy efficiency, and computational precision. However, the pursuit of higher performance necessitates more complex circuit designs and increased operating frequencies, which exacerbate IR-drop issues. Severe IR-drop can significantly degrade chip performance and even threaten reliability. Conventional circuit-level IR-drop mitigation methods, such as back-end optimizations, are resource-intensive and often compromise power, performance, and area (PPA). To address these challenges, we propose AIM, comprehensive software and hardware co-design for architecture-level IR-drop mitigation in high-performance PIM. Initially, leveraging the bit-serial and in-situ dataflow processing properties of PIM, we introduce Rtog and HR, which establish a direct correlation between PIM workloads and IR-drop. Building on this foundation, we propose LHR and WDS, enabling extensive exploration of architecture-level IR-drop mitigation while maintaining computational accuracy through software optimization. Subsequently, we develop IR-Booster, a dynamic adjustment mechanism that integrates software-level HR information with hardware-based IR-drop monitoring to adapt the V-f pairs of the PIM macro, achieving enhanced energy efficiency and performance. Finally, we propose the HR-aware task mapping method, bridging software and hardware designs to achieve optimal improvement. Post-layout simulation results on a 7nm 256-TOPS PIM chip demonstrate that AIM achieves up to 69.2% IR-drop mitigation, resulting in 2.29x energy efficiency improvement and 1.152x speedup.

存算一体电压降软硬件协同能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。