用软硬件协同设计实现超低功耗模拟循环计算,突破噪声累积瓶颈。
Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations

- 提出基于双稳态存储单元的模拟循环网络,通过离散输出抑制噪声20倍以上。
- 在180nm CMOS上实现电路与软件精确匹配,验证了线性边际成本添加循环结构。
- 适用于环境传感器、生物植入等始终在线的超低功耗场景。
始终在线的AI应用,如环境传感器和生物植入设备,需要极低功耗。模拟电路可实现亚微瓦级推理,但现有模拟实现仅限前馈架构:将其实现为循环动态被认为不切实际,因时间反馈导致噪声累积。本文通过软硬件协同设计证明该障碍可被克服。具体而言,我们发现具有离散输出和滞后动力学的双稳态存储循环单元(BMRUs)可实现超低功耗电流模式模拟实现,从头构建电路。该电路使每个学习参数与电路元件一一对应。离散输出在每单元边界至少抑制20倍模拟噪声,打破阻碍模拟循环的噪声累积。我们将BMRUs重构为第一象限固定阈值操作,实现直接对应的同时保持表达能力与可训练性。180 nm CMOS晶体管级仿真显示,软件预测与电路行为高度一致,软件模型可低成本高保真模拟物理硬件。利用此保真度,开展大规模抗噪性与功耗缩放分析:添加循环的功耗随状态维度线性增长,而主导总功耗的前馈层呈二次增长,表明循环结构以线性边际成本加入。端到端关键词识别在循环核心实现亚微瓦推理。
原文摘要 · Abstract (English)
Always-on AI applications, from environmental sensors to biomedical implants, require ultra-low power consumption. Analog circuits offer a path to sub-microwatt inference, yet existing analog implementations are limited to feedforward architectures: extending them to recurrent dynamics has been considered impractical due to noise accumulation through temporal feedback. We demonstrate that this barrier can be overcome through hardware-software co-design. Specifically, we identify that Bistable Memory Recurrent Units (BMRUs), a class of Recurrent Neural Networks (RNNs) with discrete-valued outputs and hysteretic dynamics, admit an ultra-low power current-mode analog implementation which we design from first principles. The resulting circuit establishes a one-to-one correspondence between each learned parameter and a circuit element. The discrete outputs suppress analog noise by at least 20-fold at each cell boundary, breaking the noise accumulation that prevents analog recurrence. We reformulate BMRUs for first-quadrant operation with fixed thresholds, enabling the direct correspondence while preserving expressivity and trainability. Transistor-level simulations in 180 nm Complementary Metal-Oxide-Semiconductor (CMOS) show near-perfect agreement between software predictions and circuit-level behavior, with the software model thereby serving as a high-fidelity simulator of the physical hardware at low computational cost. We leverage this fidelity to conduct large-scale noise immunity and power scaling analyses: the power cost of adding recurrence scales linearly with state dimension, while the feedforward layers dominating total power scale quadratically, meaning recurrence is added at linear marginal cost relative to the feedforward backbone. End-to-end keyword spotting achieves sub-microwatt inference at the RNN core.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。