arXiv:2508.16151cs.ARcs.CL2025-08被引 3

用金属布线固化大模型权重,实现超高效低耗推理芯片。

Hardwired-Neurons Language Processing Units as General-Purpose Cognitive Substrates

  • 将模型权重嵌入金属布线3D拓扑,替代传统硅单元阵列。
  • 每秒处理24.9万词元,能效比提升千倍,碳足迹降低357倍。
  • 适合大规模语言模型部署,尤其关注成本与能耗的场景。

大语言模型(LLM)的快速发展使语言成为通用认知基础,推动专用语言处理单元(LPUs)的需求。为应对LLM推理的高能耗问题,本文提出硬连线神经元语言处理单元(HNLPU),通过物理固化模型权重实现计算效率的数个数量级提升。然而,现代大模型规模带来巨大制造成本:直接固化gpt-oss 120B需超过60亿美元的光掩模费用,经济上不可行。为此,我们提出新型金属嵌入(Metal-Embedding)方法——将权重参数嵌入金属导线的三维拓扑结构中,实现密度提升15倍,并使70个光掩模层中有60个在芯片间保持一致,包括全部EUV掩模。整体使光掩模成本降低112倍,将非重复工程(NRE)成本降至可接受范围。实验表明,HNLPU达249,960 tokens/s(GPU/WSE的5,555倍/85倍)、36 tokens/J(1,047倍/283倍),芯片总面积13,232 mm²,5nm工艺下预估NRE为5946万至1.235亿美元。分析显示,在每年更新权重假设下,相比OpenAI规模的H100集群,HNLPU成本效益提升41.7–80.4倍,碳足迹减少357倍。

原文摘要 · Abstract (English)

The rapid advancement of Large Language Models (LLMs) has established language as a core general-purpose cognitive substrate, driving the demand for specialized Language Processing Units (LPUs) tailored for LLM inference. To overcome the growing energy consumption of LLM inference systems, this paper proposes a Hardwired-Neurons Language Processing Unit (HNLPU), which physically hardwires LLM weight parameters into the computational fabric, achieving several orders of magnitude computational efficiency improvement by extreme specialization. However, a significant challenge still lies in the scale of modern LLMs. A straightforward hardwiring of gpt-oss 120 B would require fabricating photomask sets valued at over 6 billion dollars, rendering this straightforward solution economically impractical. Addressing this challenge, we propose the novel Metal-Embedding methodology. Instead of embedding weights in a 2D grid of silicon device cells, Metal-Embedding embeds weight parameters into the 3D topology of metal wires. This brings two benefits: (1) a 15x increase in density, and (2) 60 out of 70 photomask layers are homogeneous across chips, including all EUV photomasks. In total, Metal-Embedding reduced the photomask cost by 112x, bringing the Non-Recurring Engineering (NRE) cost of HNLPU into an economically viable range. Experimental results show that HNLPU achieved 249,960 tokens/s (5,555x/85x that of GPU/WSE), 36 tokens/J (1,047x/283x that of GPU/WSE), 13,232 mm2 total die area, $59.46 M-123.5 M estimated NRE at 5 nm technology. Analysis shows that HNLPU achieved 41.7-80.4x improvement in cost-effectiveness and 357x reduction in carbon footprint compared to OpenAI-scale H100 clusters, under an annual weight updating assumption.

硬件加速大模型推理能效优化芯片设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。