在粒子对撞机中实现实时机器学习推理,需优化算法部署于FPGA硬件。
Analysis of Hardware Synthesis Strategies for Machine Learning in Collider Trigger and Data Acquisition
- 对比SNL与hls4ml框架在FPGA上的资源占用与延迟表现
- 不同模型规模下,两者在资源消耗与响应速度上各有优劣
- 为对撞机实时数据处理提供硬件选型参考,适合高能物理工程师
为充分挖掘当前及未来高能粒子对撞机的物理潜力,可在探测器电子系统中引入机器学习(ML)实现智能数据处理与采集。然而,在对撞机环境中实现真正的实时性,软件方法无法满足极低延迟要求,必须将机器学习算法优化并合成到硬件中部署。本文分析了神经网络推理效率,聚焦于场可编程门阵列(FPGA)上对撞机触发算法的应用。评估了两个框架——SLAC神经网络库(SNL)与hls4ml——在不同模型规模下的资源消耗与延迟表现,揭示了各自的优缺点,为实时、资源受限环境中的神经网络部署提供了重要指导。本工作旨在帮助研究人员和工程师选择最适合的软硬件配置。
原文摘要 · Abstract (English)
To fully exploit the physics potential of current and future high energy particle colliders, machine learning (ML) can be implemented in detector electronics for intelligent data processing and acquisition. The implementation of ML in real-time at colliders requires very low latencies that are unachievable with a software-based approach, requiring optimization and synthesis of ML algorithms for deployment on hardware. An analysis of neural network inference efficiency is presented, focusing on the application of collider trigger algorithms in field programmable gate arrays (FPGAs). Trade-offs are evaluated between two frameworks, the SLAC Neural Network Library (SNL) and hls4ml, in terms of resources and latency for different model sizes. Results highlight the strengths and limitations of each approach, offering valuable insights for optimizing real-time neural network deployments at colliders. This work aims to guide researchers and engineers in selecting the most suitable hardware and software configurations for real-time, resource-constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。