arXiv:2606.12966cs.LGcs.NE2026-06被引 2

发现模型泛化前先出现频率同步,且可提前1700步预测。

Circuit Synchronization Precedes Generalization: A Causal Precursor to Grokking

论文配图:Circuit Synchronization Precedes Generalization: A Causal Precursor to Grokking
图 1 · 摘自论文原文
  • 提出频谱同步度(FSD)量化傅里叶电路同步,无需先验知识
  • FSD在9种配置下均提前1722步达峰值,早于泛化发生
  • 同步是泛化的因果前提,调节权重衰减可控制泛化时机

Grokking 是指变换器在模算术任务上训练时,准确率从随机水平突然跃升至接近完美的延迟泛化现象。已有研究将其归因于基于傅里叶的算法电路,但其时间规律、因果结构与可操控性仍不清楚。本文引入频谱同步度(FSD),一种无需先验知识的归一化、置换检验指标,用于衡量傅里叶电路的同步程度。在九种模加法配置(五种素数,三种种子)中,FSD在平均1722步(范围500–3000)前达到后泛化水平,且在所有情况下均早于受限逻辑损失基线,成为最早的可用预测信号。我们提供了直接因果证据:在FSD峰值点分叉训练并调整权重衰减λ,可单调提前泛化时间,Δt与1/λ成正比。该规律在三个素数上复现(种子平均R²为0.89至0.99)。泛化发生在近似恒定的记忆范数上,支持阈值机制。该现象非傅里叶检测器对傅里叶电路的误判:在非阿贝尔群S5上,基于基忠实的FSD变体在六种子上均提前于泛化,而原版则不然。利用FSD峰值调度权重衰减增长,可在不破坏训练稳定性的前提下加速泛化。仅注意力模块的模型表现出强FSD前兆并实现泛化,而仅MLP模型则从未泛化。

原文摘要 · Abstract (English)

Grokking is the delayed generalisation phenomenon where a transformer trained on modular arithmetic abruptly transitions from near-chance to near-perfect validation accuracy. It has been attributed to a Fourier-based algorithmic circuit, but its timing, causal structure, and controllability remain poorly understood. We introduce the Frequency Synchronization Degree (FSD), a normalised, permutation-tested metric for Fourier circuit synchronisation requiring no prior knowledge of the circuit. Across nine modular addition configurations (five primes, three seeds), FSD reaches its post-grokking level 500 to 3000 steps before grokking (mean lead 1722 steps, every configuration positive, sign-test p approx 0.004), and synchronises before a restricted-logit loss baseline in all nine cases, making it the earliest available predictor. We give direct causal evidence that the inter-phase gap is a regularisation phenomenon: forking training at the FSD-ceiling step and varying weight decay lambda produces monotonically earlier grokking, with delta-t proportional to 1/lambda. This law replicates across three primes (R-squared 0.89 to 0.99 on seed-averaged delta-t); per-run R-squared is unstable due to the chaotic transition, so we report error bars rather than single runs. Grokking occurs at a near-constant memorisation norm across lambda, grounding the constant in a threshold mechanism. This is not an artefact of applying a Fourier detector to a Fourier circuit: on the non-abelian group S5, a basis-faithful generalisation of FSD precedes grokking on all six seeds, while the original Fourier FSD does not. Using the FSD ceiling to schedule a weight-decay increase also accelerates grokking over a fixed schedule without destabilising training. An attention-only variant groks with a strong FSD precursor while an MLP-only model never groks.

深度学习泛化机制因果分析模型行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。