arXiv:2605.11847cs.ETcs.LG2026-05

用新型锁存器结构提升类脑存储器能效,突破传统设计瓶颈。

A Fast and Energy-Efficient Latch-Based Memristive Analog Content-Addressable Memory

论文配图:A Fast and Energy-Efficient Latch-Based Memristive Analog Content-Addressable Memory
图 1 · 摘自论文原文
  • 用动态电流竞争比较替代静态分压,实现高再生增益和零静态功耗。
  • 读取能耗降低33%,在相同延迟下比6T2M架构更节能且可扩展。
  • 适合边缘AI与嵌入式智能场景,尤其适合高维数据决策树推理。

基于忆阻器的模拟内容寻址存储器(aCAM)为边缘AI与嵌入式智能应用中的大规模关联计算提供了高效能路径,已成功用于决策树推理,并拓展了存算一体(CIM)架构的功能边界。然而,传统6T2M结构存在静态搜索功耗高、电压增益有限及匹配线串扰显著等问题,制约了模拟精度与可扩展性。本文提出强臂锁存忆阻器(SALM)aCAM单元,以动态电流竞争比较器替代静态电压分压,实现高再生增益、内在结果锁存及近乎零的静态搜索功耗。相比6T2M,SALM在相同延迟下读取能耗降低33%,并消除增益与串扰限制,支持大规模阵列扩展。SALM还支持可扩展的串行与并行锁存共享机制,结合数据集感知优化框架,揭示显式的能效-延迟权衡,在代表性工作负载上实现高达50%的能耗降低(代价为3倍延迟)。为支持架构探索,我们基于22 nm FD-SOI工艺的SPICE查表构建了电路级行为模型,精确捕捉匹配线动态与串扰特性。集成至X-TIME决策树编译器后,SALM在高维数据集上保持近软件精度,而基线设计因增益受限与累积串扰导致性能下降。

原文摘要 · Abstract (English)

Analog content-addressable memories (aCAMs) based on memristors provide a promising pathway toward energy-efficient large-scale associative computing for Edge AI and embedded intelligence applications. They have been successfully applied to decision-tree inference and extend the capabilities of compute-in-memory (CIM) architectures beyond conventional vector-matrix multiplication. However, conventional designs such as the 6T2M architecture suffer from static search power, limited voltage gain, and pronounced match-line crosstalk, constraining analog precision and scalability. We introduce a strong-arm latched memristor (SALM) aCAM cell that replaces static voltage division with a dynamic current-race comparator, enabling high regenerative gain, intrinsic result latching, and near-zero static search power. Compared to 6T2M, SALM reduces read energy by 33% at identical latency while eliminating the gain and crosstalk limitations that prevent 6T2M from scaling to large arrays. SALM further enables scalable sequential and parallel latch sharing, and a dataset-aware optimization framework exposes an explicit energy-latency tradeoff, achieving up to 50% energy reduction at 3x latency across representative workloads. To enable architectural exploration, we develop a circuit-accurate behavioral model derived from SPICE lookup tables in 22 nm FD-SOI technology, capturing match-line dynamics and crosstalk. Integrated into the X-TIME decision-tree compiler, this framework demonstrates that SALM maintains near-software accuracy for high-dimensional datasets, whereas baseline designs degrade due to limited gain and cumulative crosstalk.

忆阻器存内计算能效优化边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。