揭示生成模型学规则的时机与边界,解释为何会‘抄作业’或‘创新’。
The two clocks and the innovation window: When and how generative models learn rules

- 用双时钟模型追踪规则学习与记忆复制的时机差异。
- 规则复杂度越高,模型越难学规则;数据越多,越容易复制样本。
- 提出‘创新窗口’概念,适用于扩散与自回归模型,指导模型设计。
在有限数据上训练的生成模型面临根本矛盾:其得分匹配或下一个词目标收敛于经验分布而非我们希望学习的总体分布。通过规则有效的合成任务,我们追踪了两个训练时间尺度上的表现:τ_rule(生成首次符合规则的步骤)和τ_mem(模型开始复现训练样本的步骤)。聚焦奇偶性规则并扩展至其他二元规则和组合谜题,我们刻画了这两个时钟τ_rule与τ_mem如何依赖于学习设置的关键因素。具体而言,τ_rule随规则复杂度增加而增大,随模型容量提升而减小;而τ_mem几乎不随规则变化,且与数据集大小N近似线性相关。我们定义‘创新窗口’为区间[τ_rule, τ_mem]。该窗口随数据量N增加而变宽,随规则复杂度增加而变窄,当τ_rule ≥ τ_mem时可能完全消失。这一双时钟结构同时存在于扩散模型(DiT)与自回归模型(GPT)中,仅存在架构相关的偏移。解构DiT模型的所学得分发现,优化景观随时间演变:规则有效样本的吸引盆地在τ_rule附近显著扩张,而训练样本的吸引盆地则在τ_mem附近开始主导。这些结果共同提供了一个统一且可预测的框架,解释生成模型何时以及如何表现出真正的创新。
原文摘要 · Abstract (English)
Generative models trained on finite data face a fundamental tension: their score-matching or next-token objective converges to the empirical training distribution rather than the population distribution we seek to learn. Using rule-valid synthetic tasks, we trace this tension across two training timescales: $τ_{\mathrm{rule}}$, the step at which generations first become rule-valid, and $τ_{\mathrm{mem}}$, the step at which models begin reproducing training samples. Focusing on parity and extending to other binary rules and combinatorial puzzles, we characterize how these two clocks, $τ_{\mathrm{rule}}$ and $τ_{\mathrm{mem}}$, depend on key aspects of the learning setup. Specifically, we show that $τ_{\mathrm{rule}}$ increases with rule complexity and decreases with model capacity, while $τ_{\mathrm{mem}}$ is approximately invariant to the rule and scales nearly linearly with dataset size $N$. We define the \emph{innovation window} as the interval $[τ_{\mathrm{rule}}, τ_{\mathrm{mem}}]$. This window widens with increasing $N$ and narrows with rule complexity, and may vanish entirely when $τ_{\mathrm{rule}} \geq τ_{\mathrm{mem}}$. The same two-clock structure arises in both diffusion (DiT) and autoregressive (GPT) models, with architecture-dependent offsets. Dissecting the learned score of DiT models reveals a corresponding evolution of the optimization landscapes, where rule-valid samples' basins expand substantially around $τ_{\mathrm{rule}}$, while training samples' basins begin to dominate around $τ_{\mathrm{mem}}$. Together, these results yield a unified and predictive account of when and how generative models exhibit genuine innovation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。