无需乘法器的TCN加速器,实现边缘端低功耗少样本持续学习。
Chameleon: A Multiplier-Free Temporal Convolutional Network Accelerator for End-to-End Few-Shot and Continual Learning from Sequential Data
- 统一架构支持少样本、持续学习与推理,面积开销仅0.5%。
- 在16kHz原始音频上实现端到端少样本/持续学习,准确率超96.8%。
- 40nm芯片实测功耗仅3.1μW,适合极致边缘设备使用。
边缘端设备上的本地学习可实现低延迟、隐私保护的个性化,并提升长期鲁棒性与降低维护成本。然而,从真实序列数据中以少量样本进行可扩展、低功耗的端到端片上学习仍是开放挑战。现有支持反向传播的加速器为提升学习性能牺牲了推理效率,而简化学习算法常无法达到可接受的准确率。本文提出Chameleon,通过三大贡献解决此问题:(i) 统一学习与推理架构,支持少样本学习(FSL)、持续学习(CL)和推理,仅增加0.5%面积开销;(ii) 利用时序卷积网络(TCNs)高效捕捉长时依赖,首次实现端到端片上少样本与持续学习及16kHz原始音频推理;(iii) 双模式无乘法器计算阵列,可匹配当前最先进仅推理关键词识别(KWS)加速器的功耗,或实现4.3倍更高峰值GOPs。40nm CMOS工艺流片验证,其在Omniglot数据集上端到端片上少样本学习准确率达96.8%(5类1样本)、98.8%(5类5样本),持续学习最终准确率为82.2%(学习250类,每类10样本),同时在12类Google语音命令数据集上保持93.3%推理准确率,功耗仅为3.1μW。
原文摘要 · Abstract (English)
On-device learning at the edge enables low-latency, private personalization with improved long-term robustness and reduced maintenance costs. Yet, achieving scalable, low-power end-to-end on-chip learning, especially from real-world sequential data with a limited number of examples, is an open challenge. Indeed, accelerators supporting error backpropagation optimize for learning performance at the expense of inference efficiency, while simplified learning algorithms often fail to reach acceptable accuracy targets. In this work, we present Chameleon, leveraging three key contributions to solve these challenges. (i) A unified learning and inference architecture supports few-shot learning (FSL), continual learning (CL) and inference at only 0.5% area overhead to the inference logic. (ii) Long temporal dependencies are efficiently captured with temporal convolutional networks (TCNs), enabling the first demonstration of end-to-end on-chip FSL and CL on sequential data and inference on 16-kHz raw audio. (iii) A dual-mode, multiplier-free compute array allows either matching the power consumption of state-of-the-art inference-only keyword spotting (KWS) accelerators or enabling $4.3\times$ higher peak GOPS. Fabricated in 40-nm CMOS, Chameleon sets new accuracy records on Omniglot for end-to-end on-chip FSL (96.8%, 5-way 1-shot, 98.8%, 5-way 5-shot) and CL (82.2% final accuracy for learning 250 classes with 10 shots), while maintaining an inference accuracy of 93.3% on the 12-class Google Speech Commands dataset at an extreme-edge power budget of 3.1 $μ$W.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。