arXiv:2601.04270cs.LG2026-01

发现深度学习梯度有可预测的低秩结构,可用两个新指标量化训练复杂度。

Predictable Gradient Manifolds in Deep Learning: Temporal Path-Length and Intrinsic Rank as a Complexity Regime

  • 用可计算的路径长度和可预测秩衡量梯度演化规律
  • 实验证明各类模型梯度轨迹均具局部可预测性与强低秩特征
  • 适合研究优化动态、设计自适应算法的研究者阅读

深度学习优化过程存在未被最坏情况梯度界捕捉的内在结构。实验表明,训练轨迹上的梯度往往具有时间可预测性,并在低维子空间中演化。本文通过可度量的框架形式化这一现象,引入两个可计算量:基于预测的路径长度(衡量梯度从历史信息中可预测的程度),以及可预测秩(量化梯度增量的内在时间维度)。我们证明经典在线与非凸优化的收敛性与遗憾边界可重新表述为依赖于这些量,而非最坏情况变化。在卷积网络、视觉变压器、语言模型及合成控制任务中,均发现梯度轨迹具有局部可预测性与强低秩结构。这些性质在不同架构与优化器间保持稳定,且可通过轻量级随机投影直接从记录梯度诊断。结果为理解现代深度学习优化动态提供了统一视角,将标准训练重定义为处于低复杂度时间区间。该视角提示了自适应优化器、秩感知追踪与基于真实训练运行可测特性的预测算法设计的新方向。

原文摘要 · Abstract (English)

Deep learning optimization exhibits structure that is not captured by worst-case gradient bounds. Empirically, gradients along training trajectories are often temporally predictable and evolve within a low-dimensional subspace. In this work we formalize this observation through a measurable framework for predictable gradient manifolds. We introduce two computable quantities: a prediction-based path length that measures how well gradients can be forecast from past information, and a predictable rank that quantifies the intrinsic temporal dimension of gradient increments. We show how classical online and nonconvex optimization guarantees can be restated so that convergence and regret depend explicitly on these quantities, rather than on worst-case variation. Across convolutional networks, vision transformers, language models, and synthetic control tasks, we find that gradient trajectories are locally predictable and exhibit strong low-rank structure over time. These properties are stable across architectures and optimizers, and can be diagnosed directly from logged gradients using lightweight random projections. Our results provide a unifying lens for understanding optimization dynamics in modern deep learning, reframing standard training as operating in a low-complexity temporal regime. This perspective suggests new directions for adaptive optimizers, rank-aware tracking, and prediction-based algorithm design grounded in measurable properties of real training runs.

优化动态梯度结构低秩分析可预测性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。