arXiv:2511.01872cs.DCcs.LG2025-11

用学习模型更准预测可重构硬件上模型部署的吞吐量

Learned Cost Model for Placement on Reconfigurable Dataflow Hardware

  • 基于数据驱动方法学习映射策略的吞吐量,替代人工设计模型
  • 相比传统方法预测准确率提升31%-52%,且无需性能标注
  • 能直接加速编译后模型,最高提升5.6%运行速度

将机器学习模型的数据流图映射到可重构系统上极具挑战性,因为不同映射方式的吞吐量差异大,且资源消耗各不相同。为解决此问题,需建立一种评估映射吞吐量的模型,而完全测量吞吐量成本过高。现有方法多依赖人工设计的解析模型,依赖代理特征或直觉,易引入误差。本文提出一种学习型方法,可在多种数据流图上将吞吐量预测准确率提升31%-52%。此外,该方法在移除性能标注后仍保持准确度,无下降。实验表明,使用该模型可使编译后的图运行速度快5.6%。

原文摘要 · Abstract (English)

Mapping a dataflow-graph of an ML model onto a reconfigurable system is difficult, as different mappings have different throughputs and consume resource constraints differently. To solve this, a model to evaluate the throughput of mappings is necessary as measuring throughput completely is expensive. Many use a hand-designed analytical model, relying on proxy features or intuition, introducing error. We provide a Learned Approach that predicts throughput 31%-52% more accurately over a variety of graphs. In addition, our approach shows no accuracy degradation after removing performance annotations. We show that using this approach results in 5.6% faster compiled graphs.

硬件优化模型部署学习模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。