arXiv:2602.18523cs.LGcs.AI2026-02被引 8

多任务模型突现泛化时,权重衰减调控着解的几何结构与稳定性。

The Geometry of Multi-Task Grokking: Transverse Instability, Superposition, and Weight Decay Phase Structure

  • 通过梯度正交性揭示优化路径约束在低维流形上。
  • 乘法任务最先泛化,且不同随机种子结果一致,呈现顺序性。
  • 权重衰减调节泛化时间,过低则完全失效,存在临界相变点。

Grokking——在接近零训练损失后突然从记忆转向泛化——此前主要研究集中于单任务设置。本文将几何分析扩展至多任务模运算问题,使用共享主干的Transformer在双任务(模加+模乘)和三任务(模加+模乘+模平方)目标下系统地进行权重衰减扫描。五种一致现象浮现:(1) 阶段性泛化顺序:乘法最先泛化,随后是平方,最后是加法,各种子结果一致;(2) 普遍可积性:优化轨迹始终被限制在经验上不变的低维执行流形内;对易子缺陷若垂直于该流形,则可靠预示泛化发生;(3) 权重衰减相结构:泛化时间尺度、曲率深度、重构阈值与缺陷领先量随权重衰减系统共变,揭示不同动力学区域及无衰减下的尖锐失败模式;(4) 全息不可压缩性:最终解仅占据4–8个主轨迹方向,却分布于满秩权重中,微小扰动即破坏性能;奇异值截断、幅度剪枝与均匀缩放均无法保持性能;(5) 横向脆弱性与冗余性:移除不足10%的正交梯度分量即可消除grokking,但双任务模型在极端删除下仍部分恢复,表明过参数化支持冗余中心流形。这些结果共同支持一种动态图像:多任务grokking在参数空间构建紧凑的叠加子空间,权重衰减作为压缩压力,而过剩参数提供优化路径的几何冗余。

原文摘要 · Abstract (English)

Grokking -- the abrupt transition from memorization to generalization long after near-zero training loss -- has been studied mainly in single-task settings. We extend geometric analysis to multi-task modular arithmetic, training shared-trunk Transformers on dual-task (mod-add + mod-mul) and tri-task (mod-add + mod-mul + mod-sq) objectives across a systematic weight decay sweep. Five consistent phenomena emerge. (1) Staggered grokking order: multiplication generalizes first, followed by squaring, then addition, with consistent delays across seeds. (2) Universal integrability: optimization trajectories remain confined to an empirically invariant low-dimensional execution manifold; commutator defects orthogonal to this manifold reliably precede generalization. (3) Weight decay phase structure: grokking timescale, curvature depth, reconstruction threshold, and defect lead covary systematically with weight decay, revealing distinct dynamical regimes and a sharp no-decay failure mode. (4) Holographic incompressibility: final solutions occupy only 4--8 principal trajectory directions yet are distributed across full-rank weights and destroyed by minimal perturbations; SVD truncation, magnitude pruning, and uniform scaling all fail to preserve performance. (5) Transverse fragility and redundancy: removing less than 10% of orthogonal gradient components eliminates grokking, yet dual-task models exhibit partial recovery under extreme deletion, suggesting redundant center manifolds enabled by overparameterization. Together, these results support a dynamical picture in which multi-task grokking constructs a compact superposition subspace in parameter space, with weight decay acting as compression pressure and excess parameters supplying geometric redundancy in optimization pathways.

多任务学习Grokking几何优化权重衰减

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。