40nm芯片实现低延迟端侧小样本学习,能效比领先
FSL-HDnn: A 40 nm Few-shot On-Device Learning Accelerator with Integrated Feature Extraction and Hyperdimensional Computing
- 用权值聚类+超维计算,免梯度更新,单次通过训练
- 40nm芯片上实现6毫焦/图能效,28图/秒吞吐
- 适合资源受限的边缘设备实时小样本学习
本文提出FSL-HDnn,一种基于40 nm CMOS工艺的能效型加速器,实现了特征提取与端侧小样本学习(FSL)的端到端流程。该加速器通过两个协同模块解决资源受限边缘场景下的端侧学习(ODL)难题:一是采用权值聚类的参数高效特征提取器以降低计算复杂度;二是基于超维计算(HDC)的小样本分类器,避免梯度反向传播,支持单次通过训练并显著降低延迟。此外,FSL-HDnn通过两项优化策略实现低延迟ODL与推理:包含分支特征提取的早退机制,以及提升硬件利用率的批量单次通过训练。实测结果表明,该芯片在10类5样本任务下达到6 mJ/image的优异训练能效和28 images/s的端到端训练吞吐率,相比现有最优ODL芯片,训练延迟降低2倍至20.9倍。
原文摘要 · Abstract (English)
This paper introduces FSL-HDnn, an energy-efficient accelerator that implements the end-to-end pipeline of feature extraction and on-device few-shot learning (FSL). The accelerator addresses fundamental challenges of on-device learning (ODL) for resource-constrained edge applications through two synergistic modules: a parameter-efficient feature extractor employing weight clustering and an FSL classifier based on hyperdimensional computing (HDC). The feature extractor exploits the weight clustering mechanism to reduce computational complexity, while the HDC-based FSL classifier eliminates gradient-based back propagation operations, enabling single-pass training with substantially reduced latency. Additionally, FSL-HDnn enables low-latency ODL and inference via two proposed optimization strategies, including an early-exit mechanism with branch feature extraction and batched single-pass training that improves hardware utilization. Measurement results demonstrate that our chip fabricated in a 40 nm CMOS process delivers superior training energy efficiency of 6 mJ/image and end-to-end training throughput of 28 images/s on a 10-way 5-shot FSL task. The end-to-end training latency is also reduced by 2x to 20.9x compared to state-of-the-art ODL chips.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。