arXiv:2608.02991cs.LG2026-08

通过可控的权重与偏置更新分配,提升神经网络优化的精度与稳定性。

Control Allocation in Neural Network Optimization: Joint Affine Control of Weight and Bias Updates

  • 引入坐标预条件指数和谱指数,实现对参数更新分布的独立控制。
  • 适度控制可提升验证集与最差组表现,过度控制则导致欠拟合。
  • 适用于需要精细调节训练过程的研究者或高性能优化场景。

优化算法不仅决定神经网络更新的大小,还决定更新在各参数通道间的分布。本文研究这种分布是否可作为独立于全局训练进度的可控量。通过归一化通道能量定义操作性更新分配,分析两个标量控制:坐标预条件指数和缩放增广权重-偏置矩阵偏置列的仿射谱指数。在固定状态下,相同的非零步长乘子保持归一化分配不变;坐标指数产生显式逆的仿射成对对数几率;仿射指数引发秩一半正定格拉姆扰动和逻辑回归原始参与规律。进一步分离原始仿射参与、谱增益与解码后的物理偏置更新,并证明有限多项式谱迭代保持奇异子空间不变。同状态重播验证了精确控制律。五种子控制基准实验显示,中间控制提升保留集与最差组指标,过度仿射控制导致欠拟合。四任务单种子迁移研究提供描述性佐证。结果确立瞬时分配控制与有界经验操作区间,但不意味着任务无关的泛化排序。

原文摘要 · Abstract (English)

Optimization algorithms determine not only the magnitude of a neural-network update but also how that update is distributed across parameter channels. We study whether this distribution can be treated as a controllable quantity independently of global training progress. We define operational update allocation through normalized channel energies and analyze two scalar controls: a coordinate-preconditioning exponent and an affine spectral exponent that scales the bias column of an augmented weight--bias matrix. At a frozen state, a common nonzero step-size multiplier leaves normalized allocation unchanged; the coordinate exponent yields affine pairwise log-odds with an explicit inverse; and the affine exponent induces a rank-one positive-semidefinite Gram perturbation and a logistic raw-participation law. We further separate raw affine participation, spectral gain, and the decoded physical bias update, and show that finite polynomial spectral iterations preserve singular subspaces. Same-state replay verifies the exact control laws. On a five-seed controlled benchmark, intermediate controls improve held-out and worst-group metrics, whereas excessive affine control causes underfitting. A four-task single-seed transfer study provides descriptive corroboration. These results establish instantaneous allocation control and a bounded empirical operating regime, but do not imply a task-independent generalization ordering.

优化算法神经网络参数控制训练稳定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。