arXiv:2608.28295cs.AI2026-08

用结构化正交算子替代密集矩阵,实现高效低功耗的类忆阻储层计算。

Memristive-Friendly Hadamard Reservoir Computing: Structured, Multiplier-Free Recurrences at Scale

论文配图:Memristive-Friendly Hadamard Reservoir Computing: Structured, Multiplier-Free Recurrences at Scale
图 1 · 摘自论文原文
  • 采用符号对角线、置换和快速沃尔什-哈达玛变换构建无乘法器的正交算子。
  • 在8192节点规模下性能媲美密集正交储层,硬件加速达50倍且内存减少10000倍。
  • 适合部署于类忆阻硬件,对噪声、量化误差等具有可预测鲁棒性。

储层计算(RC)通过固定不变的递归层设计循环神经网络,是类忆阻硬件的理想选择。现有忆阻友好型储层虽基于忆阻器件动力学模拟神经元行为,但仍依赖密集递归矩阵,物理实现成本高昂。本文提出用结构化正交算子替代密集矩阵,该算子由符号对角线、置换和快速沃尔什-哈达玛变换构成,无需乘法器,每步仅需$O(N)$参数和$O(N\log N)$操作,且不显式存储矩阵。我们在标准与忆阻友好型回声状态网络中实现了该算子,每个单元仅需一个二值输入连接。数学分析表明,精确正交性使回声态条件在递归尺度上紧致,噪声响应可在设计阶段预知。实验在20个分类与7个回归基准上验证,最大规模达$N=8192$,结构化模型性能媲美密集正交储层,平均表现优于环形储层,且优势随规模增大而扩大。在三种硬件平台上的实测显示,递归步骤速度最高提升50倍,内存占用降低$10^4$倍。最后我们进行了算子消融实验,量化了噪声、量化、器件失配及离散故障的影响。

原文摘要 · Abstract (English)

Reservoir Computing (RC) designs Recurrent Neural Networks around a fixed, i.e., untrained, recurrent layer, and is a natural candidate for neuromorphic hardware. Memristive-friendly reservoirs derive the neuron dynamics from memristive-device kinetics, but still rely on dense recurrent matrices, which are expensive to realize physically. In this paper, we replace the dense matrix with a structured orthogonal operator, built from sign diagonals, a permutation, and a fast Walsh-Hadamard transform. The operator is multiplier-free, requires $O(N)$ parameters and $O(N\log N)$ operations per step, and is never materialized as a matrix. We instantiate it in a standard and in a memristive-friendly Echo State Network, with one binary input connection per unit. Our mathematical analysis shows that exact orthogonality yields an echo state condition that is tight in the recurrent scaling, and a noise response that is predictable at design time. Moreover, the operator mixes the whole state in a single application. Experiments on twenty classification and seven regression benchmarks, at reservoir sizes up to $N = 8192$, show that the structured models match dense orthogonal reservoirs, and achieve better mean performance than the cycle reservoir by a margin that widens with size. Furthermore, we time the recurrent step on three hardware platforms, where it is up to $50\times$ faster than a dense product and $10^4\times$ smaller in memory. Finally, we ablate the operator and measure the response to noise, quantization, device mismatch and discrete faults.

储层计算忆阻器正交算子低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。