arXiv:2604.23806cs.LGcs.AI2026-04

用物理硬件实现扩散模型训练,能耗降低千倍以上。

Symmetric Equilibrium Propagation for Thermodynamic Diffusion Training

论文配图:Symmetric Equilibrium Propagation for Thermodynamic Diffusion Training
图 1 · 摘自论文原文
  • 直接在模拟硬件上用对称平衡传播训练,无需数字加速器
  • 实测每步训练能耗比GPU低1000到10000倍
  • 适合追求极致能效的神经网络物理实现研究者

基于分数的扩散模型的反向过程在形式上等价于时变能量景观下的过阻尼朗之万动力学。我们此前的工作表明,通过用低秩模块间耦合替代密集跳跃连接,双线性耦合类比硬件可在投影层面实现该动力学,相比数字推理具有三至四个数量级的能效优势。然而,该硬件能否闭合训练环路——即不依赖外部数字加速器传输梯度——仍悬而未决。本文给出肯定回答:将平衡传播直接应用于双线性能量函数,在零扰动极限下可得到无偏的去噪分数匹配梯度估计。对于有限扰动,我们推导出仅由硬件刚度、局部曲率及损失梯度范数控制的精确偏差界,并发现耦合参数更新存在一个主导偏差项恒为零的双线性特例。对称扰动进一步将主要偏差从 $ \mathcal{O}(β) $ 提升至 $ \mathcal{O}(β^2) $,代价极低。在真实有限松弛预算下,单边平衡传播产生反相关梯度,而对称版本则输出对齐更新。偏差-方差分析确定了最优工作点,端到端物理单元核算显示,每训练步相比匹配的GPU基准有 $ 10^3$-$10^4\times $ 的能效优势。对称双线性平衡传播是首个保持低秩耦合、仅读出即可完成的局部训练规则,使可扩展热力学扩散模型成为可能。

原文摘要 · Abstract (English)

The reverse process in score-based diffusion models is formally equivalent to overdamped Langevin dynamics in a time-dependent energy landscape. In our prior work we showed that a bilinearly-coupled analog substrate can physically realize this dynamics at a projected three-to-four orders of magnitude energy advantage over digital inference by replacing dense skip connections with low-rank inter-module couplings. Whether the \emph{training} loop can be closed on the same substrate -- without routing gradients through an external digital accelerator -- has remained open. We resolve this affirmatively: Equilibrium Propagation applied directly to the bilinear energy yields an unbiased estimator of the denoising score-matching gradient in the zero-nudge limit. For finite nudging we derive a sharp bias bound controlled solely by substrate stiffness, local curvature, and the norm of the loss-gradient signal, with a bilinear-specific corollary showing that one dominant bias term vanishes identically for coupling-parameter updates. Symmetric nudging further upgrades the leading bias from $ \mathcal{O}(β) $ to $ \mathcal{O}(β^2) $ at negligible extra cost. Under realistic finite-relaxation budgets this upgrade is essential, as one-sided EqProp produces anti-correlated gradients while symmetric EqProp yields well-aligned updates. Bias-variance analysis determines the optimal operating point, and end-to-end physical-unit accounting projects a $ 10^3$-$10^4\times $ energy advantage per training step over a matched GPU baseline. Symmetric bilinear EqProp is the first local, readout-only training rule that preserves the low-rank coupling enabling scalable thermodynamic diffusion models.

扩散模型物理计算能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。