提出首个支持全六种微标度数据类型的可扩展硬件,提升机器人边缘持续学习效率。
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
- 采用子字并行与统一整浮点处理,实现六种微标度数据类型兼容。
- 通过共享指数分组减少存储冗余,训练吞吐量提升4倍,内存占用降低51%。
- 适合需要低功耗持续学习的机器人边缘计算场景。
自主机器人需在本地高效学习以适应新环境,避免依赖云端。微标度(MX)数据类型通过结合整数与浮点表示并共享指数,可在保持精度的同时降低能耗。然而,现有连续学习处理器Dacapo仅支持MXINT且反向传播中向量分组效率低下。本文首次提出两项创新:(1) 精度可扩展算术单元,利用子字并行与统一整浮点处理,支持全部六种MX数据类型;(2) 支持平方共享指数分组,实现反向传播中权重高效处理,消除存储冗余与量化开销。我们在TSMC 16nm FinFET工艺、400MHz下,于四个机器人工作负载上对齐峰值吞吐量评估该设计,相比Dacapo实现51%更小内存占用、4倍更高有效训练吞吐量,同时保持相近能效,为边缘端机器人持续学习提供高效支持。
原文摘要 · Abstract (English)
Autonomous robots require efficient on-device learning to adapt to new environments without cloud dependency. For this edge training, Microscaling (MX) data types offer a promising solution by combining integer and floating-point representations with shared exponents, reducing energy consumption while maintaining accuracy. However, the state-of-the-art continuous learning processor, namely Dacapo, faces limitations with its MXINT-only support and inefficient vector-based grouping during backpropagation. In this paper, we present, to the best of our knowledge, the first work that addresses these limitations with two key innovations: (1) a precision-scalable arithmetic unit that supports all six MX data types by exploiting sub-word parallelism and unified integer and floating-point processing; and (2) support for square shared exponent groups to enable efficient weight handling during backpropagation, removing storage redundancy and quantization overhead. We evaluate our design against Dacapo under iso-peak-throughput on four robotics workloads in TSMC 16nm FinFET technology at 400MHz, reaching a 51% lower memory footprint, and 4x higher effective training throughput, while achieving comparable energy efficiency, enabling efficient robotics continual learning at the edge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。