arXiv:2510.01578cs.LG2025-10中稿 · as a conference pa…被引 3

提出平滑的梯度调控方法,让训练更稳定高效。

Gradient Shaping Beyond Clipping: A Functional Perspective on Update Magnitude Control

  • 基于层内梯度统计动态调整更新幅度,替代固定阈值裁剪
  • 在图像与语言任务中提升收敛速度与训练稳定性
  • 适合追求训练鲁棒性与优化效率的研究者

梯度裁剪广泛用于稳定深度网络训练,但其作为硬性固定阈值的设定限制了灵活性,并忽略梯度分布动态。本文提出 SPAMP(统计逐层自适应调制与投影),将裁剪泛化为平滑的逐层梯度调控框架。SPAMP 跟踪局部梯度统计,动态估计阈值,并通过幂变换对更新幅度进行可微调制。这一视角将裁剪与预热视为控制有效更新尺度 $η_t \|g_t\|$ 的互补机制,提供了比僵化启发式更合理的替代方案。在图像与语言任务上的大量实验表明,SPAMP 在稳定性、收敛性和鲁棒性方面优于现有方法。

原文摘要 · Abstract (English)

Gradient clipping is widely used to stabilize deep network training, but its formulation as a hard, fixed threshold limits flexibility and ignores gradient distribution dynamics. We propose SPAMP (Statistical Per-layer Adaptive Modulation and Projection), a unified framework that generalizes clipping into smooth, per-layer gradient shaping. SPAMP tracks local gradient statistics, dynamically estimates thresholds, and applies power-based transformations to modulate update magnitudes in a differentiable manner. This perspective recasts clipping and warmup as dual mechanisms for controlling the effective update scale $η_t \|g_t\|$, offering a principled alternative to rigid heuristics. Extensive experiments across image and language tasks demonstrate that SPAMP improves stability, convergence, and robustness over existing methods.

梯度调控训练稳定可微调制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。