arXiv:2607.18574cs.LGcs.NE2026-07

通过活动与误差几何条件化,显著提升直接反馈对齐的训练效果。

Conditioned Direct Feedback Alignment via Activity and Error Geometry

  • 分离活动与误差因素,分别进行条件化优化
  • 活动条件化可提升近40个百分点,误差条件化增益1.77–7.53个百分点
  • 适用于非卷积模型,尤其在高噪声任务中表现突出

直接反馈对齐(DFA)通过固定随机投影输出误差来训练隐藏层,避免了反向传播(BP)中的转置权重反传。我们研究了DFA训练中一种不同于反馈质量的失效模式:局部权重更新由外积计算,其各向异性可能源于前突触活动或局部误差因子。在受控合成场景中,我们发现当高方差方向包含任务无关干扰时,活动条件化可带来约40个百分点的性能提升。三次清晰验证揭示另一情形:误差条件化使原始DFA提升1.77–7.53个百分点,独立选择活动与误差因子组合后额外提升0.40–0.90个百分点。该结果在tanh/one-vs-rest MNIST和预注册Fashion-MNIST上成立,并在ReLU/softmax MNIST模型中八次新种子实验复现。该分解构建了对称块局部归一化DFA(nDFA)家族:活动nDFA右预条件化逆活动二阶矩,误差nDFA左预条件化逆局部误差二阶矩,K-nDFA同时应用双因子并独立调参阻尼。线性化后对齐计算给出精确输入侧谱恒等式,支持双侧规则的克罗内克因子动机;范数匹配排除标量步长解释。误差因子在欠阻尼时敏感,BatchNorm是有效的活动侧替代方案,卷积网络增益仍有限。因此,本工作将条件化DFA定位为对局部外积规则失效机制的研究,而非对BP的通用替代或全层卷积信用分配的解法。

原文摘要 · Abstract (English)

Direct feedback alignment (DFA) trains hidden layers with fixed random projections of the output error, avoiding the transposed-weight backward pass of backpropagation (BP). We study a failure mode of DFA training that is distinct from feedback quality: the local weight update is calculated by an outer product, so anisotropy can enter through either its presynaptic-activity factor or its local-error factor. Our analyses with controlled synthetic regimes isolate the first failure mode and show an approximately 40-percentage-point activity-conditioning gain when high-variance directions contain task-irrelevant nuisance. Three clean confirmations isolate a different regime: error conditioning improves raw DFA by 1.77--7.53 percentage points, and combining independently selected activity and error factors adds 0.40--0.90 points over activity conditioning. The signs hold for tanh/one-vs-rest MNIST and preregistered Fashion-MNIST, and replicate on eight fresh seeds in a ReLU/softmax MNIST model. This factorization yields a symmetric block-local family of normalized DFA (nDFA): activity nDFA right-preconditions by an inverse activity second moment, error nDFA left-preconditions by an inverse local-error second moment, and K-nDFA applies both factors with separately tuned damping. A linearized post-alignment calculation gives an exact input-side spectral identity and a Kronecker-factor motivation for the two-sided rule, whereas norm matching rules out a scalar step-size explanation. The error factor is fragile when under-damped, BatchNorm is a strong activity-side alternative, and convnet gains remain partial. We therefore frame conditioned DFA as a factor-level study of when local outer-product rules fail, not as a general replacement for BP or a solution to all-layer convolutional credit assignment.

神经网络反馈对齐机器学习优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。