arXiv:2601.22690cs.LGcs.AI2026-01

探究Transformer在周期性模式泛化上的能力边界

Do Transformers Have the Ability for Periodicity Generalization?

  • 从抽象代数视角统一解释单/复合周期性规律
  • 模型仅能记忆训练数据,无法泛化到未见的复合周期
  • 构建可控生成基准Coper,支持两种分布外测试场景

基于Transformer的大语言模型在多种任务中表现优异,但在分布外(OOD)泛化能力上仍远逊于人类。本文以周期性这一基本OOD场景为切入点,研究其背后机制。周期性体现为变化中的不变性,周期性泛化指模型从训练数据中提取周期模式并推广至未知场景的能力。我们从抽象代数角度统一阐释了单周期与复合周期,并揭示了Transformer在此类任务中的局限性。为此,我们构建了名为Coper的可控生成基准,包含两种分布外设置:Hollow与Extrapolation。实验表明,尽管模型可在训练中记忆周期数据,但无法泛化至未见过的复合周期结构。代码已开源,供后续研究使用。

原文摘要 · Abstract (English)

Large language models (LLMs) based on the Transformer have demonstrated strong performance across diverse tasks. However, current models still exhibit substantial limitations in out-of-distribution (OOD) generalization compared with humans. We investigate this gap through periodicity, one of the basic OOD scenarios. Periodicity captures invariance amid variation. Periodicity generalization represents a model's ability to extract periodic patterns from training data and generalize to OOD scenarios. We introduce a unified interpretation of periodicity from the perspective of abstract algebra and reasoning, including both single and composite periodicity, to explain why Transformers struggle to generalize periodicity. Then we construct Coper about composite periodicity, a controllable generative benchmark with two OOD settings, Hollow and Extrapolation. Experiments reveal that periodicity generalization in Transformers is limited, where models can memorize periodic data during training, but cannot generalize to unseen composite periodicity. We release the source code to support future research.

周期性泛化TransformerOOD泛化生成基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。