arXiv:2604.13543cs.ARcs.LG2026-04中稿 · IEEE ISQED'26

为实时步态分析设计了跨层协同优化的LSTM加速器,提升效率并降低功耗。

Cross-Layer Co-Optimized LSTM Accelerator for Real-Time Gait Analysis

  • 从软件到版图全程优化,通过量化减少硬件复杂度
  • 在65nm工艺下实现0.325mm²芯片面积,提速4.05倍于需求
  • 适合边缘计算场景的高精度、低功耗步态异常检测应用

长短期记忆(LSTM)神经网络已广泛应用于医疗领域,其中实时性与边缘计算能力至关重要。步态分析通过检测异常步态以预防跌倒,是典型应用场景。鉴于性能、功耗和面积的严苛要求,专用集成电路(ASIC)能高效实现LSTM的实时部署,保持高准确率。本文首次提出面向实时步态分析的跨层协同优化LSTM加速器,针对ASIC设计进行系统级探索。从软件层开展硬件感知的位宽优化以降低硬件复杂度,于寄存器传输级探索多种架构,并生成不同版图方案以权衡硬件复杂度与准确率。物理综合结果表明,在65 nm工艺下,高精度优化版布局面积为0.325 mm²;另一兼顾硬件复杂度的替代设计面积缩小15.4%,且整体加速器处理速度较应用需求快4.05倍。

原文摘要 · Abstract (English)

Long Short-Term Memory (LSTM) neural networks have penetrated healthcare applications where real-time requirements and edge computing capabilities are essential. Gait analysis that detects abnormal steps to prevent patients from falling is a prominent problem for such applications. Given the extremely stringent design requirements in performance, power dissipation, and area, an Application-Specific Integrated Circuit (ASIC) enables an efficient real-time exploitation of LSTMs for gait analysis, achieving high accuracy. To the best of our knowledge, this work presents the first cross-layer co-optimized LSTM accelerator for real-time gait analysis, targeting an ASIC design. We conduct a comprehensive design space exploration from software down to layout design. We carry out a bit-width optimization at the software level with hardware-aware quantization to reduce the hardware complexity, explore various designs at the register-transfer level, and generate alternative layouts to find efficient realizations of the LSTM accelerator in terms of hardware complexity and accuracy. The physical synthesis results show that, using the 65 nm technology, the die size of the accelerator's layout optimized for the highest accuracy is 0.325 mm^2, while the alternative design optimized for hardware complexity with a slightly lower accuracy occupies 15.4% smaller area. Moreover, the designed accelerators achieve accurate gait abnormality detection 4.05x faster than the given application requirement.

LSTM加速边缘计算步态分析ASIC设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。