arXiv:2507.01003cs.LGcs.AI2025-07

用李雅普诺夫指数诊断训练状态,引入幽灵节点加速早期训练。

Description of the Training Process of Neural Networks via Ergodic Theorem : Ghost nodes

  • 通过运行时最大李雅普诺夫指数判断是否真正收敛。
  • 添加幽灵输出节点,绕过狭窄损失屏障提升初期训练效率。
  • 训练后期幽灵节点自动消失,不影响最终模型性能。

近期研究从遍历性视角解释神经网络训练过程。本文在此基础上提出统一框架,通过分析目标函数的几何结构,引入运行时最大李雅普诺夫指数的估计值,可严格区分真实收敛至稳定极小值与仅在鞍点附近统计稳定的现象。进而提出标准分类器的幽灵节点扩展:增加辅助幽灵输出节点,使模型获得额外下降方向,在早期训练阶段形成横向通道,避开低质量损失盆地。我们证明该扩展严格降低近似误差;充分收敛后,幽灵维度会坍缩,扩展模型恢复为原模型,且在增广参数空间中存在一条总损失不增加的路径。整体结果提供了一种可解释、架构友好的干预策略,既加速早期训练,又保持渐近行为,同时兼具正则化作用。

原文摘要 · Abstract (English)

Recent studies have proposed interpreting the training process from an ergodic perspective. Building on this foundation, we present a unified framework for understanding and accelerating the training of deep neural networks via stochastic gradient descent (SGD). By analyzing the geometric landscape of the objective function we introduce a practical diagnostic, the running estimate of the largest Lyapunov exponent, which provably distinguishes genuine convergence toward stable minimizers from mere statistical stabilization near saddle points. We then propose a ghost category extension for standard classifiers that adds auxiliary ghost output nodes so the model gains extra descent directions that open a lateral corridor around narrow loss barriers and enable the optimizer to bypass poor basins during the early training phase. We show that this extension strictly reduces the approximation error and that after sufficient convergence the ghost dimensions collapse so that the extended model coincides with the original one and there exists a path in the enlarged parameter space along which the total loss does not increase. Taken together, these results provide a principled architecture level intervention that accelerates early stage trainability while preserving asymptotic behavior and simultaneously serves as an architecture-friendly regularizer.

神经网络训练优化方法深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。