arXiv:2607.12501cs.LGeess.IV2026-07

提出一种新型训练方法,解决前向-前向算法的尺度失控问题。

Gauge-Fixing the Forward-Forward Objective: A Whitened Goodness Derived from a Likelihood-Ratio Account

  • 引入白化不变的得分函数,避免权重缩放带来的误差
  • 在9个实验组合中,线性探测准确率提升4至7点
  • 适合关注无反向传播训练的模型研究者

前向-前向算法通过局部训练每一层,使真实输入的激活平方和得分高,对比样本得分低。在显式生成模型下,该得分是似然比检验的充分统计量,而成对目标存在规范自由度:可通过放大权重降低损失,而非分离数据分布。本文提出修复方案——在线层内训练的白化、尺度不变得分,经评估表明,在三个语料库、三种深度和四倍层宽范围(每组13次种子)下,所有测量单元均优于标准成对目标,9组中有8组提升4-7个百分点,缩小了与端到端反向传播差距的16%-61%。对照实验显示,辛顿固定阈值损失虽能限制失控(1.4倍对比133倍),但无法实现不变性,仅靠边界控制不具优势。相较于现有最优的稀疏顶k得分,新方法在两个语料库上表现相当,但唯一彻底消除尺度失控。所有边界与预测均在实验前记录,反驳结果也已报告。

原文摘要 · Abstract (English)

The Forward-Forward algorithm trains each layer locally, so that a scalar goodness - the sum of squared activations - is high on real inputs and low on contrastive ones. Under an explicit generative model this goodness is the sufficient statistic of a likelihood-ratio test, and the pairwise form of the objective admits a gauge: a layer can lower its loss by inflating the scale of its weights rather than by separating the two populations. The analysis prescribes the repair - a whitened, scale-invariant goodness trained online within each layer - which we evaluate as a training procedure. Across three corpora, three depths and a fourfold range of layer width (13 seeds per cell), it raises linear-probe accuracy over the standard pairwise objective in every measured cell - by 4 to 7 points on eight of nine corpus-depth combinations - and closes 16-61% of the gap to end-to-end backpropagation. A control isolates the mechanism: Hinton's fixed-threshold loss also bounds the runaway, to a factor of 1.4 against 133, yet tracks the unmodified baseline - invariance to the gauge, not a bound on it, is what pays. Against the strongest published alternative - a sparse, top-k goodness - the derived objective is statistically indistinguishable on two corpora of three, yet only it removes the runaway: sparsity and gauge-invariance are independent axes, and the published variant recovers accuracy while leaving the pathology in place. We state the boundaries we measured, and every prediction was recorded before its experiment with the refutations reported.

前向-前向无反向传播模型稳定性白化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。