arXiv:2608.15483cs.LG2026-08

通过预测性分析揭示神经网络训练中的时间结构规律。

Measuring Structured Predictability in Neural Training Dynamics: A Cross-Regime Study

论文配图:Measuring Structured Predictability in Neural Training Dynamics: A Cross-Regime Study
图 1 · 摘自论文原文
  • 设计三类探测器,量化参数更新的短期可预测性。
  • 辅助参数动态简单,主参数预测性集中在局部时变区域。
  • 适合研究训练过程动力学的模型开发者和算法设计师。

现代深度网络通过长时间更新轨迹进行训练,但其时间组织机制仍缺乏系统刻画。本文以短期可预测性为度量,探究近期更新中是否蕴含未来参数变化的信息。结合位移方向、子空间残差和基于预测器的三类探测器,采用符合惯例的零基准校准组级读出方法,应用于在CIFAR上的多轮视觉训练及公开的Pythia预训练检查点。结果表明:如归一化参数和偏置等向量型张量(辅助参数)表现出更简单的短期动态,而矩阵型特征变换权重(主参数)的可预测行为集中于局部且随时间变化的“口袋”中。各类探测器间的一致性及其与独立轨迹诊断的一致性表明,测量结果捕捉到了内在轨迹结构;探测器差异则揭示了不同形式的时间组织。在CIFAR上的控制实验进一步显示,架构与训练方案系统性地影响可测结构。对Pythia-70M的案例研究揭示了一系列分层、尺度与角色依赖的事件,包括主参数的ESA低于随机符号一致水平,以及qkv可预测性“口袋”在各层间的出现与重分布。这些结果将短期可预测性定位为一种可回溯、参数级的训练动态诊断工具。

原文摘要 · Abstract (English)

Modern deep networks are trained through long update trajectories, yet their temporal organization remains less systematically characterized than architectures, losses, or optimizers. We study short-horizon predictability as a measure of temporal redundancy: where, when, and under which training conditions recent updates contain information about near-future parameter motion. We combine three complementary probe families, displacement-direction, subspace-residual, and predictor-based probes, with convention-aware, null-calibrated group-level readouts, and apply them to multi-pass vision training on CIFAR and public Pythia pretraining checkpoints. Across both regimes, vector-like tensors such as normalization parameters and biases (auxiliary parameters) exhibit simpler short-horizon dynamics than matrix-like feature-transforming weights (bulk parameters), whose predictable behavior concentrates in localized, time-varying pockets. Agreement within and across probe families, and with independent trajectory diagnostics, indicates that these measurements capture intrinsic trajectory structure, while probe differences distinguish complementary forms of temporal organization. Controlled CIFAR comparisons further show that architecture and training recipe systematically modulate the measured structure. A Pythia-70M case study further exposes a sequence of role-, depth-, and scale-dependent events, including bulk ESA falling below the random sign-agreement level and the emergence and redistribution of predictable qkv pockets across layers. These results position short-horizon predictability as a retrospective, parameter-resolved diagnostic of training dynamics.

训练动态可预测性神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。