arXiv:2509.22410cs.ARcs.LG2025-09被引 1

用深度学习实现真实应用下的快速精准处理器性能预测

NeuroScalar: A Deep Learning Framework for Fast, Accurate, and In-the-Wild Cycle-Level Performance Prediction

  • 基于微架构无关特征训练模型,预测新处理器的周期级性能
  • 在消费级GPU上达5兆指令/秒,性能开销仅0.1%
  • 可部署于现有芯片,支持大规模真实应用的硬件对比测试

新处理器设计的评估受限于速度慢、依赖非代表性基准测试的周期级仿真器。本文提出一种全新的深度学习框架,可在真实生产环境中实现高保真度的‘在野’仿真。核心贡献是训练一个基于微架构无关特征的深度学习模型,用于预测假设性处理器设计的周期级性能。这一独特方法使模型可直接部署于现有硅片上,以评估未来硬件。我们构建了一个完整的系统,包含轻量级硬件追踪采集器和有原则的采样策略,最大限度减少用户影响。该系统在通用GPU上实现5兆指令/秒的仿真速度,性能开销仅为0.1%。此外,协同设计的Neutrino片上加速器相比GPU性能提升85倍。实验证明,该框架可在大规模真实应用中实现准确的性能分析与海量硬件A/B测试。

原文摘要 · Abstract (English)

The evaluation of new microprocessor designs is constrained by slow, cycle-accurate simulators that rely on unrepresentative benchmark traces. This paper introduces a novel deep learning framework for high-fidelity, ``in-the-wild'' simulation on production hardware. Our core contribution is a DL model trained on microarchitecture-independent features to predict cycle-level performance for hypothetical processor designs. This unique approach allows the model to be deployed on existing silicon to evaluate future hardware. We propose a complete system featuring a lightweight hardware trace collector and a principled sampling strategy to minimize user impact. This system achieves a simulation speed of 5 MIPS on a commodity GPU, imposing a mere 0.1% performance overhead. Furthermore, our co-designed Neutrino on-chip accelerator improves performance by 85x over the GPU. We demonstrate that this framework enables accurate performance analysis and large-scale hardware A/B testing on a massive scale using real-world applications.

性能预测深度学习硬件加速真实应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。