arXiv:2504.19797cs.ARcs.LG2025-04被引 12

用可重构的逻辑电路实现边缘设备上的高效训练,比深度网络省电6倍。

Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs

  • 基于有限状态机在FPGA上实现动态可重构训练
  • 相比同类设计能效高2.54倍,功耗低6倍
  • 适合多源传感器数据的边缘实时学习任务

物联网节点对数据隐私与安全的需求推动了边缘训练的发展。传统深度神经网络因反向传播复杂、精度权衡和层异构性,在资源受限的边缘设备上部署困难。本文提出一种动态刘特曼机(DTM)训练加速器,基于门控逻辑与有限状态自动机实现片上学习,可在同一FPGA芯片内完成推理与训练。该设计支持运行时重构,适配不同数据集、模型架构与规模,无需重新综合。与深度神经网络相比,DTM训练减少乘加操作,无需导数计算;其数据驱动的学习机制通过对齐刘特曼自动机与输入数据形成逻辑命题,实现高效的查找表映射和低内存占用。所提加速器在能效方面达2.54倍于现有最佳方案,功耗降低6倍。

原文摘要 · Abstract (English)

The increased demand for data privacy and security in machine learning (ML) applications has put impetus on effective edge training on Internet-of-Things (IoT) nodes. Edge training aims to leverage speed, energy efficiency and adaptability within the resource constraints of the nodes. Deploying and training Deep Neural Networks (DNNs)-based models at the edge, although accurate, posit significant challenges from the back-propagation algorithm's complexity, bit precision trade-offs, and heterogeneity of DNN layers. This paper presents a Dynamic Tsetlin Machine (DTM) training accelerator as an alternative to DNN implementations. DTM utilizes logic-based on-chip inference with finite-state automata-driven learning within the same Field Programmable Gate Array (FPGA) package. Underpinned on the Vanilla and Coalesced Tsetlin Machine algorithms, the dynamic aspect of the accelerator design allows for a run-time reconfiguration targeting different datasets, model architectures, and model sizes without resynthesis. This makes the DTM suitable for targeting multivariate sensor-based edge tasks. Compared to DNNs, DTM trains with fewer multiply-accumulates, devoid of derivative computation. It is a data-centric ML algorithm that learns by aligning Tsetlin automata with input data to form logical propositions enabling efficient Look-up-Table (LUT) mapping and frugal Block RAM usage in FPGA training implementations. The proposed accelerator offers 2.54x more Giga operations per second per Watt (GOP/s per W) and uses 6x less power than the next-best comparable design.

边缘计算FPGA加速低功耗逻辑学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。