模拟电路系统通过特定训练也能实现双下降,突破传统数字模型限制。
Analog Physical Systems Can Exhibit Double Descent
- 用自调整电阻网络构建无数字处理器的模拟系统
- 修正训练协议后成功实现双下降现象
- 为类脑计算和生物系统提供新思路
大型人工智能模型成功的关键之一是双下降现象:当模型规模超过数据量时,性能反而提升。本文首次在去中心化的模拟电阻网络中实现了这一现象。该系统无需数字处理器即可自我训练并完成任务,具备能耗低、速度快的潜力,但需应对元件非理想性问题。我们发现标准训练无法实现双下降,而改进后的训练协议可克服此类缺陷。结果表明,只要采用适当训练方法,模拟物理系统也能表现出支撑数字AI成功的机制。研究还暗示生物系统可能同样受益于过参数化。
原文摘要 · Abstract (English)
An important component of the success of large AI models is double descent, in which networks avoid overfitting as they grow relative to the amount of training data, instead improving their performance on unseen data. Here we demonstrate double descent in a decentralized analog network of self-adjusting resistive elements. This system trains itself and performs tasks without a digital processor, offering potential gains in energy efficiency and speed -- but must endure component non-idealities. We find that standard training fails to yield double descent, but a modified protocol that accommodates this inherent imperfection succeeds. Our findings show that analog physical systems, if appropriately trained, can exhibit behaviors underlying the success of digital AI. Further, they suggest that biological systems might similarly benefit from over-parameterization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。