ReLU神经网络训练中,梯度下降与连续流的导数不等价,导致敏感性分析偏差。
Singular Curvature in ReLU Training:Differentiation and the Gradient-Flow Limit Need Not Commute
- 通过精确离散导数分析,揭示了梯度下降与连续梯度流在导数上的本质差异
- 在有限时域内,离散导数收敛到无事件区域传播器,而连续流导数包含速度归一化激活事件转移
- 适用于全批量、确定性、有限步长的稳定训练场景,对敏感性建模有重要启示
梯度下降(GD)是梯度流的显式欧拉近似,但状态准确的连续时间代理在微分后未必保持准确。在任意固定的非共振步长下,普通自动微分可精确计算执行的硬ReLU GD程序的导数。我们证明,在固定有限时域内,GD状态收敛,这些精确离散导数趋近于无事件区域传播器,而极限流的导数则包含速度归一化的激活事件转移。预点斯特尔茨表示将绝对连续的区域海森矩阵与原子界面曲率分离;一个非零梯度跳跃产生精确秩一的端点偏差,且只要事件严格存在,全局凸性就无法实现多事件完全抵消。然而,一类全局1-强凸的残差ReLU平方损失风险在开放初始化集上可实现任意大的倒敏感比,具有统一横截裕度。相同的离散-流分解也适用于参数和反向模式伴随量;在标量或自治归一化情形下解析平滑及一致事件定位可恢复流敏感性。结果针对确定性的全批量、有限时域动力学,具有稳定有限路径的分离同向横截事件;它们是一致性定理,而非大规模训练中的普遍性声明。
原文摘要 · Abstract (English)
Gradient descent (GD) is explicit Euler for gradient flow, but a state-accurate continuous-time surrogate need not remain accurate after differentiation. At every fixed nonresonant step size, ordinary automatic differentiation exactly differentiates the executed hard-ReLU GD program. We prove that, over a fixed finite horizon, the GD states converge and these exact discrete derivatives approach an event-free regional propagator, whereas the derivative of the limiting flow also contains speed-normalized activation-event transfers. A prepoint Stieltjes representation separates the absolutely continuous regional Hessian from atomic interface curvature; one nonzero gradient jump produces an exactly rank-one endpoint discrepancy, and global convexity prevents complete multi-event cancellation whenever an event is strict. Nevertheless, a standard family of globally 1-strongly convex residual-ReLU squared-loss risks realizes arbitrarily large reciprocal sensitivity ratios on open initialization sets, with a uniform transversality margin. The same discrete-versus-flow decomposition extends to parameters and reverse-mode adjoints; resolved smoothing in the scalar or autonomous-normal regime and consistent event localization recover the flow sensitivity. The results concern deterministic full-batch, finite-horizon dynamics with a stable finite itinerary of separated same-direction transverse events; they are consistency theorems, not prevalence claims for large-scale training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。