预测小维度随机重参数训练神经网络的可行条件
Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks
- 基于曲率谱与位移分布,提出定向解析公式预测训练成功率
- 实验证明该方法在图像和语言模型中精准捕捉训练临界点
- 新结构可大幅降低存储与内存开销,适合高效微调场景
神经网络可通过随机低维重参数化进行训练或微调,即用一个小型潜在向量通过固定的随机映射生成完整的参数更新。这引出关键问题:潜在空间需多大才能达到低损失区域?本文将已知的可达性相变转化为以紧凑凸目标为中心的锥形形式,基于极锥的统计维数。主要理论贡献是提出一个定向解析的二次主公式,可同时利用曲率谱和参考解位移轨迹预测随机切片残差。该公式导出自洽的各向同性定向预测器,并在仅考虑半径的保守特例下恢复早期高斯宽度二次界。基于此分析,我们提出随机映射网络(RaMaN),采用结构化的Hadamard或种子再生高斯映射实现预测的潜维数。此类构造避免了密集随机映射的O(dP)存储开销,将优化器状态内存从O(P)降至O(d)。同时开发无矩阵曲率近似与免扫掠维度选择。在可控二次型与神经曲率实验中,定向解析预测器紧密跟踪实测过渡位置,在位移方向重要时优于各向同性近似。端到端实验进一步显示,图像与语言模型中存在显著且依赖协议的训练突变。
原文摘要 · Abstract (English)
Neural networks can often be trained or fine-tuned through random low-dimensional reparameterization, where a small latent vector is mapped into a full parameter update by a frozen random map. This raises a practical question: how large must the latent search space be to reach a low-loss region? We first express the known accessibility transition in an equivalent conic form, centered for compact convex targets at the statistical dimension of the polar cone. Our main theoretical contribution is an orientation-resolved quadratic master formula that predicts the random-slice residual from both the curvature spectrum and the reference-to-solution displacement profile. It yields a self-consistent isotropic-orientation predictor and, in a conservative radius-only specialization, recovers the earlier Gaussian-width quadratic bound. Building on this analysis, we introduce Random Mapping Networks (RaMaN), which instantiate the predicted latent dimension using structured Hadamard or seed-regenerated Gaussian maps. These constructions avoid the O(dP) storage of dense random maps and reduce optimizer-state memory from O(P) to O(d). We also develop matrix-free curvature approximations and sweep-free dimension selection. Across controlled quadratic and neural-curvature experiments, the orientation-resolved predictor closely tracks measured transition locations and outperforms orientation-agnostic approximations when displacement direction matters. End-to-end experiments further show sharp, protocol-dependent training transitions across image and language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。