arXiv:2606.20869cs.ARcs.AI2026-06

AI算法与硬件协同设计,自动生成高效模型-加速器组合。

A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generation

论文配图:A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generation
图 1 · 摘自论文原文
  • 联合搜索算法与硬件设计空间,打破传统分步开发流程。
  • 自动生成兼顾精度、延迟、能效的模型-加速器配对方案。
  • 适合需要极致性能优化的嵌入式AI系统研发人员。

我们提出一种面向人工智能算法与加速器协同设计、协同搜索和协同生成(A3C3)的完整方法论,通过联合优化神经网络架构及其硬件实现,解决传统自上而下设计流程中的效率低下问题。传统AI部署常将模型设计与硬件映射分阶段进行:先追求精度设计算法,再适配延迟、吞吐、能耗或资源约束。这种分离导致系统性能不佳,尤其在现代异构、内存密集且平台依赖的AI工作负载下更为明显。A3C3则参数化算法与加速器的设计空间,并进行联合搜索,实现模型-加速器对的自动生成,更好地平衡精度、延迟、吞吐、能效与硬件利用率。本文为Sudeep Pasricha和Muhammad Shafique主编的《嵌入式机器学习手册》(Springer Nature)的章节。

原文摘要 · Abstract (English)

We present a holistic methodology for artificial intelligence algorithm and accelerator co-design, co-search, and co-generation (A3C3), which jointly optimizes neural network architectures and their hardware implementations to address the inefficiencies of traditional top-down AI system design flows. Conventional AI deployment often treats model design and hardware mapping as separate stages: an algorithm is first developed for accuracy, and only afterward adapted to meet latency, throughput, energy, or resource constraints. This separation can lead to suboptimal systems, particularly as modern AI workloads become increasingly heterogeneous, memory-intensive, and platform-dependent. A3C3 instead parameterizes both algorithmic and accelerator design spaces and searches them jointly, enabling the automatic generation of model-accelerator pairs that better balance accuracy, latency, throughput, energy efficiency, and hardware utilization. This article is a book chapter of the Handbook of Embedded Machine Learning, edited by Sudeep Pasricha and Muhammad Shafique, Springer Nature.

协同设计AI加速器自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。