arXiv:2604.27911cs.LGcs.ET2026-04被引 3

用物理硬件直接实现大模型,让AI更省电、更快、更小。

Physical Foundation Models: Fixed hardware implementations of large-scale neural networks

论文配图:Physical Foundation Models: Fixed hardware implementations of large-scale neural networks
图 1 · 摘自论文原文
  • 把神经网络直接做成物理结构,利用材料自然特性运行计算。
  • 10^12参数模型可大幅降低数据中心能耗,支持边缘设备部署。
  • 未来或可实现10^15甚至10^18参数的超大规模模型。

基础模型是训练于海量数据的深度神经网络(如GPT-5、Gemini~3、Opus~4),可执行文本生成、代码生成、问答、摘要、图像分类等多样化下游任务。其核心理念是投入资源构建单一、大规模(约10^12参数)通用模型,通过极少或无需额外训练即可适配多种任务。我们提出,基础模型的兴起为硬件工程师带来新机遇:与其为不同任务设计专用芯片,不如以约每年一次的节奏,制造固定硬件实现新一代基础模型。除传统数字电子推理硬件外,我们倡导更激进的思路——将神经网络直接体现在物理结构中,借助硬件固有物理动态进行运算,即‘物理基础模型’(PFM)。PFM有望在能效、速度和参数密度上实现数量级提升。对约10^12参数模型而言,这不仅能缓解数据中心高能耗问题,还可使当前受功耗限制的边缘设备运行更大规模模型。此外,PFM或可支持远超现有规模的模型:10^15甚至10^18参数的模型在某些度量下已具可行性。本文以光学为例(三维纳米玻璃介质)进行粗略估算,展示PFM扩展潜力,并探讨纳米电子学及其他物理平台的可能性。最后,我们总结了实现万亿参数级及以上PFM所需攻克的重大研究挑战。

原文摘要 · Abstract (English)

Foundation models are deep neural networks (such as GPT-5, Gemini~3, and Opus~4) trained on large datasets that can perform diverse downstream tasks -- text and code generation, question answering, summarization, image classification, and so on. The philosophy of foundation models is to put effort into a single, large (${\sim}10^{12}$-parameter) general-purpose model that can be adapted to many downstream tasks with no or minimal additional training. We argue that the rise of foundation models presents an opportunity for hardware engineers: in contrast to when different models were used for different tasks, it now makes sense to build special-purpose, fixed hardware implementations of neural networks, manufactured and released at the roughly 1-year cadence of major new foundation-model versions. Beyond conventional digital-electronic inference hardware with read-only weight memory, we advocate a more radical re-thinking: hardware in which the neural network is realized directly at the level of the physical design and operates via the hardware's natural physical dynamics -- \textit{Physical Foundation Models} (PFMs). PFMs could enable orders-of-magnitude advantages in energy efficiency, speed, and parameter density. For ${\sim}10^{12}$-parameter models, this would both reduce the high energy burden of AI in datacenters and enable AI in edge devices that today are power-constrained to far smaller models. PFMs could also enable inference hardware for models much larger than current ones: $10^{15}$- or even $10^{18}$-parameter PFMs seem plausible by some measures. We present back-of-the-envelope calculations illustrating PFM scaling using an optical example -- a 3D nanostructured glass medium -- and discuss prospects in nanoelectronics and other physical platforms. We conclude with the major research challenges that must be resolved for trillion-parameter PFMs and beyond to become reality.

物理计算硬件加速大模型能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。