用FPGA与AI引擎协同加速实时图神经网络,提升粒子对撞机触发系统性能。
Reconfigurable Computing Challenge: Real-Time Graph Neural Networks for Online Event Selection in Big Science

- 融合算子、分区映射与空间并行,实现端到端自动化设计流程。
- 每秒处理294万事件,延迟7.15微秒,吞吐量提升53%。
- 适合高能物理实时数据处理场景,可交互可视化推理结果。
图神经网络在对撞机实验触发系统中的应用日益广泛,但严格的延迟与吞吐量要求使嵌入式部署面临挑战。随着探测器精细化程度提高,单次推理输入数量增加,纯FPGA方案遭遇资源瓶颈。本文提出面向Belle II电磁量能器硬件触发的动态图神经网络实时部署端到端演示系统,基于AMD Versal VCK190平台,结合FPGA逻辑阵列与AI Engine核心。我们开发了基于Python的半自动化设计流程,涵盖算子融合、任务分区、映射、空间并行及核级优化。最终实现每秒294万事件的吞吐量,端到端延迟7.15微秒。相比纯FPGA基线,吞吐量提升53%,DSP资源占用从99%降至19%,AI Engine利用率仅29%。为验证部署效果,构建了交互式可视化管道,支持在真实硬件上实时监控推理结果。
原文摘要 · Abstract (English)
Graph neural networks are increasingly adopted in trigger systems for collider experiments, where strict latency and throughput constraints render deployment on embedded platforms challenging. As detectors move towards higher granularity, the number of inputs per inference increase and FPGA-only solutions face resource bottlenecks. This work presents an end-to-end demonstrator for the real-time deployment of a dynamic Graph Neural Network for the Belle II electromagnetic calorimeter hardware trigger on the AMD Versal VCK190, leveraging both FPGA fabric and AI Engine tiles. We develop a Python-based semi-automated design flow covering operator fusion, partitioning, mapping, spatial parallelization, and kernel-level optimization. Our design achieves a throughput of 2.94 million events per second at an end-to-end latency of 7.15 microseconds. Compared to the FPGA-only baseline, this represents a 53% throughput improvement while reducing DSP utilization from 99% to 19% at 29% AI Engine tile utilization. To validate the deployment, an interactive visualization pipeline enables real-time monitoring of inference results on the physical demonstrator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。