arXiv:2604.17673cs.LG2026-04

扩散模型在模加法任务中出现延迟泛化,揭示了其符号推理机制。

Grokking of Diffusion Models: Case Study on Modular Addition

论文配图:Grokking of Diffusion Models: Case Study on Modular Addition
图 1 · 摘自论文原文
  • 通过流匹配训练,模型在过拟合后突然学会模加法计算。
  • 在多样图像下,模型在关键时间步将任务拆分为算术与去噪两阶段。
  • 适合研究生成模型如何实现符号运算的学者参考。

尽管扩散模型在实践中表现优异,但其泛化机制仍不明确。我们发现,采用流匹配目标训练的扩散模型在模加法任务上表现出grokking现象——即在过拟合后出现延迟泛化,从而可对内部计算过程进行受控分析。研究覆盖两种数据设置:单图像情形下,模型通过组合操作数的周期性表征实现模加法;在高类内变异性、多图像情形下,模型利用迭代采样过程,将任务分解为算术计算阶段与视觉去噪阶段,二者由一个临界时间步分隔。本工作揭示了扩散模型在算法学习中的机械分解机制,展示了其如何连接连续像素空间生成与离散符号推理。

原文摘要 · Abstract (English)

Despite their empirical success, how diffusion models generalize remains poorly understood from a mechanistic perspective. We demonstrate that diffusion models trained with flow-matching objectives exhibit grokking--delayed generalization after overfitting--on modular addition, enabling controlled analysis of their internal computations. We study this phenomenon across two levels of data regime. In a single-image regime, mechanistic dissection reveals that the model implements modular addition by composing periodic representations of individual operands. In a diverse-image regime with high intraclass variability, we find that the model leverages its iterative sampling process to partition the task into an arithmetic computation phase followed by a visual denoising phase, separated by a critical timestep threshold. Our work provides the mechanistic decomposition of algorithmic learning in diffusion models, revealing how these models bridge continuous pixel-space generation and discrete symbolic reasoning.

扩散模型符号推理算法学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。