开源框架SNAX提升多加速器系统效率,解决软硬件兼容与数据传输难题。
An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator Systems
- 采用松耦合控制与紧耦合数据访问的混合协同机制。
- 在低功耗异构SoC上实现神经网络性能超10倍提升,加速器利用率超90%。
- 提供可定制编译器与可复用硬件模块,适合系统级开发人员使用。
异构加速器中心计算集群正成为应对多样化AI工作负载的高效方案。然而,现有集成策略常导致数据移动效率低下,并在软硬件层面存在兼容性问题,难以实现性能与易用性的统一。为此,我们提出SNAX——一个开源的软硬件协同开发框架,通过创新的混合耦合架构(松耦合异步控制 + 紧耦合数据访问),支持高效的多加速器平台构建。SNAX提供可复用硬件模块以提升计算加速器利用率,并配备基于MLIR的可定制编译器,自动完成关键系统管理任务,实现定制化多加速器集群的快速开发与部署。大量实验表明,在低功耗异构SoC上,各加速器可轻松集成与编程,相比其他加速器系统,神经网络性能提升超过10倍,且全系统运行时加速器利用率保持在90%以上。
原文摘要 · Abstract (English)
Heterogeneous accelerator-centric compute clusters are emerging as efficient solutions for diverse AI workloads. However, current integration strategies often compromise data movement efficiency and encounter compatibility issues in hardware and software. This prevents a unified approach that balances performance and ease of use. To this end, we present SNAX, an open-source integrated HW-SW framework enabling efficient multi-accelerator platforms through a novel hybrid-coupling scheme, consisting of loosely coupled asynchronous control and tightly coupled data access. SNAX brings reusable hardware modules designed to enhance compute accelerator utilization, and its customizable MLIR-based compiler to automate key system management tasks, jointly enabling rapid development and deployment of customized multi-accelerator compute clusters. Through extensive experimentation, we demonstrate SNAX's efficiency and flexibility in a low-power heterogeneous SoC. Accelerators can easily be integrated and programmed to achieve > 10x improvement in neural network performance compared to other accelerator systems while maintaining accelerator utilization of > 90% in full system operation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。