用解析与机器学习融合方法,极速高准预测芯片性能。
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
- 基于组件性能分布建模,不逐条模拟指令
- 比传统仿真快5个数量级,平均CPI误差仅2%
- 适合芯片架构设计与性能归因分析
周期级模拟器如gem5在微架构设计中广泛应用,但大规模设计空间探索时速度极慢。本文提出Concorde,一种快速且准确的微架构性能建模方法。与现有模拟和学习方法不同,Concorde不逐条模拟指令,而是基于紧凑的性能分布来预测程序行为,这些分布通过简单解析模型估算各微架构组件对性能的影响边界,从而在大量微架构参数空间中提供丰富而简洁的性能表征。实验表明,Concorde比参考周期级模拟器快超过五个数量级,跨SPEC、开源及专有基准测试的平均CPI预测误差约为2%。该速度使此前无法实现的快速设计空间探索与性能敏感性分析成为可能,例如在一小时内完成对多种程序中不同微架构组件的细粒度性能归因,共需约1.5亿次CPI评估。
原文摘要 · Abstract (English)
Cycle-level simulators such as gem5 are widely used in microarchitecture design, but they are prohibitively slow for large-scale design space explorations. We present Concorde, a new methodology for learning fast and accurate performance models of microarchitectures. Unlike existing simulators and learning approaches that emulate each instruction, Concorde predicts the behavior of a program based on compact performance distributions that capture the impact of different microarchitectural components. It derives these performance distributions using simple analytical models that estimate bounds on performance induced by each microarchitectural component, providing a simple yet rich representation of a program's performance characteristics across a large space of microarchitectural parameters. Experiments show that Concorde is more than five orders of magnitude faster than a reference cycle-level simulator, with about 2% average Cycles-Per-Instruction (CPI) prediction error across a range of SPEC, open-source, and proprietary benchmarks. This enables rapid design-space exploration and performance sensitivity analyses that are currently infeasible, e.g., in about an hour, we conducted a first-of-its-kind fine-grained performance attribution to different microarchitectural components across a diverse set of programs, requiring nearly 150 million CPI evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。