用树状张量网络实现亚微秒级低延迟机器学习推理
Ultra-low latency quantum-inspired machine learning predictors implemented on FPGA
- 基于树状张量网络设计轻量化分类器,适配FPGA硬件
- 在经典数据集与高能物理数据上实现全流水线推理
- 实测延迟低于1微秒,适合高频实时场景
张量网络(TNs)是一种用于表示量子多体系统的计算范式。近期研究表明,TNs也可用于执行机器学习任务,结果可与传统监督学习方法媲美。本文研究了树状张量网络(TTNs)在高频实时应用中的潜力,利用现场可编程门阵列(FPGA)的低延迟特性。我们实现了多种TTN分类器,可在经典机器学习数据集及复杂物理数据上进行推理。训练阶段进行了键维数、权重量化、纠缠熵与关联测量的预分析,以优化TTN架构选择。生成的TTNs被部署于硬件加速器,通过集成在服务器中的FPGA完全卸载推理任务。最终实现了一个用于高能物理(HEP)应用的分类器,以全流水线方式运行,延迟低于1微秒。
原文摘要 · Abstract (English)
Tensor Networks (TNs) are a computational paradigm used for representing quantum many-body systems. Recent works have shown how TNs can also be applied to perform Machine Learning (ML) tasks, yielding comparable results to standard supervised learning techniques. In this work, we study the use of Tree Tensor Networks (TTNs) in high-frequency real-time applications by exploiting the low-latency hardware of the Field-Programmable Gate Array (FPGA) technology. We present different implementations of TTN classifiers, capable of performing inference on classical ML datasets as well as on complex physics data. A preparatory analysis of bond dimensions and weight quantization is realized in the training phase, together with entanglement entropy and correlation measurements, that help setting the choice of the TTN architecture. The generated TTNs are then deployed on a hardware accelerator; using an FPGA integrated into a server, the inference of the TTN is completely offloaded. Eventually, a classifier for High Energy Physics (HEP) applications is implemented and executed fully pipelined with sub-microsecond latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。