arXiv:2601.14087cs.ARcs.AI2026-01中稿 · oral presentation …

用1比特计数排序单元降低DNN加速器互联功耗

'1'-bit Count-based Sorting Unit to Reduce Link Power in DNN Accelerators

  • 通过近似计算将比特计数分组,实现无比较排序
  • 面积减少35.4%,仍保持19.50%的链路功耗降低
  • 适合追求能效的CNN硬件加速器设计

深度神经网络(DNN)加速器中的互连功耗仍是瓶颈。基于1比特计数对数据排序可减少开关活动,但实际硬件排序方案仍研究不足。本文提出一种专为卷积神经网络(CNN)优化的无比较排序单元硬件实现。通过近似计算将种群计数分入粗粒度桶中,该设计在保持数据重排带来的链路功耗优势的同时显著降低硬件面积。相比精确实现的20.42%功耗降低,该近似排序单元仍维持19.50%的功耗降低,同时实现最高达35.4%的面积缩减。

原文摘要 · Abstract (English)

Interconnect power consumption remains a bottleneck in Deep Neural Network (DNN) accelerators. While ordering data based on '1'-bit counts can mitigate this via reduced switching activity, practical hardware sorting implementations remain underexplored. This work proposes the hardware implementation of a comparison-free sorting unit optimized for Convolutional Neural Networks (CNN). By leveraging approximate computing to group population counts into coarse-grained buckets, our design achieves hardware area reductions while preserving the link power benefits of data reordering. Our approximate sorting unit achieves up to 35.4% area reduction while maintaining 19.50\% BT reduction compared to 20.42% of precise implementation.

DNN加速器低功耗设计近似计算排序单元

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。