arXiv:2510.13320cs.LG2025-10被引 2

RockNet让超低功耗设备分布式训练模型,精度超现有方法2倍。

RockNet: Distributed Learning on Ultra-Low-Power Devices

  • 多设备协同训练,用高效通信协议降低计算与能耗
  • 在20个设备上从零开始训练,分类精度提升至现有方法2倍
  • 适合物联网、工业传感等边缘设备的实时故障检测场景

随着机器学习在网络物理系统中的普及,将训练从云端转移到本地设备(TinyML)成为趋势,以应对隐私和延迟问题。然而,网络物理系统常由超低功耗微控制器组成,其算力有限,难以实现训练。本文提出RockNet,一种专为超低功耗硬件设计的新型TinyML方法,在时间序列分类任务(如故障或恶意软件检测)中达到业界领先精度,无需离线预训练。通过利用系统中多个设备的协同能力,设计了一种结合机器学习与无线通信的分布式学习方法。RockNet利用所有设备并行训练轻量级且计算高效的分类器,通信开销极小。配合定制的高效多跳无线通信协议,有效克服了分布式学习中的通信瓶颈。在包含20个超低功耗设备的实验平台上验证,RockNet可从零开始成功训练时间序列分类任务,精度最高比最新微控制器神经网络训练方法高出2倍。当设备数从1增至20时,每个设备的内存、延迟和能耗降低最多达90%。结果表明,分布式机器学习、计算与通信的深度融合,首次实现了在超低功耗硬件上的高精度训练。

原文摘要 · Abstract (English)

As Machine Learning (ML) becomes integral to Cyber-Physical Systems (CPS), there is growing interest in shifting training from traditional cloud-based to on-device processing (TinyML), for example, due to privacy and latency concerns. However, CPS often comprise ultra-low-power microcontrollers, whose limited compute resources make training challenging. This paper presents RockNet, a new TinyML method tailored for ultra-low-power hardware that achieves state-of-the-art accuracy in timeseries classification, such as fault or malware detection, without requiring offline pretraining. By leveraging that CPS consist of multiple devices, we design a distributed learning method that integrates ML and wireless communication. RockNet leverages all devices for distributed training of specialized compute efficient classifiers that need minimal communication overhead for parallelization. Combined with tailored and efficient wireless multi-hop communication protocols, our approach overcomes the communication bottleneck that often occurs in distributed learning. Hardware experiments on a testbed with 20 ultra-low-power devices demonstrate RockNet's effectiveness. It successfully learns timeseries classification tasks from scratch, surpassing the accuracy of the latest approach for neural network microcontroller training by up to 2x. RockNet's distributed ML architecture reduces memory, latency and energy consumption per device by up to 90 % when scaling from one central device to 20 devices. Our results show that a tight integration of distributed ML, distributed computing, and communication enables, for the first time, training on ultra-low-power hardware with state-of-the-art accuracy.

TinyML分布式训练边缘计算低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。