arXiv:2605.27946stat.MLcs.LG2026-05

合成梯度在特定条件下可显著提升神经网络训练样本效率。

Is Backpropagation Optimal? When Synthetic Gradients Improve Sample Efficiency

论文配图:Is Backpropagation Optimal? When Synthetic Gradients Improve Sample Efficiency
图 1 · 摘自论文原文
  • 提出统一向量化反馈框架,让合成梯度成为反向传播的自然替代
  • 理论证明合成梯度在某些情况下梯度估计误差更低,优势可无限大
  • 在上下文猜谜和强化学习任务中验证了其实际潜力

反向传播是可微神经网络的标准学习规则,常被视为默认方案。本文从样本效率角度重新审视这一惯例。我们提出一个统一的向量化反馈框架,适用于基于损失和基于奖励的学习,其中合成梯度作为反向传播的自然替代出现。我们刻画了合成梯度相比反向传播实现更低梯度估计均方误差的条件,并构造了示例说明该样本效率优势可任意大。在上下文猜谜和强化学习任务上的实验验证了理论发现的潜力。

原文摘要 · Abstract (English)

Backpropagation is the default learning rule for artificial neural networks and is often treated as the settled approach whenever differentiability is available. In this work, we revisit this convention through a theoretical lens of sample efficiency. We introduce a unified vectorized feedback framework for loss-based and reward-based learning on computational graphs, in which synthetic gradients emerge as a natural alternative to backpropagation. We characterize the conditions under which synthetic gradients can achieve a lower gradient-estimation mean squared error than backpropagation. We construct examples illustrating that this sample efficiency advantage can be arbitrarily large. Experiments on contextual bandits and reinforcement learning tasks demonstrate the potential of our theoretical findings.

梯度优化强化学习神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。