arXiv:2605.28864cs.AIcs.CL2026-05

用范畴论设计新模型,让语言模型更聪明。

The Cognitive Categorical Transformer: Category-Theoretic Inductive Biases for Language Modeling

  • 引入范畴论中的单纯形消息传递机制,增强模型结构理解能力。
  • 在WikiText-103上降低2.92困惑度,相对提升12%。
  • 首次验证单纯形消息传递对3亿参数模型有效,适合结构建模研究者。

认知范畴变换器(CCT)是一个306M参数的架构,基于预训练GPT-2 Small,并融合来自范畴论和认知科学的认知基础组件。在相同训练设置(215,000次优化器步数、相同数据、相同优化器与调度)下,于WikiText-103上达到21.27的验证困惑度,相比同条件下微调的GPT-2 Small基线(24.19)降低2.92 PPL(相对减少12%)。通过从头训练消融实验,排除完整单纯形消息传递后得到23.72 PPL,表明84%(2.45/2.92)的性能提升源自GT-Full。这是首个在306M参数规模上验证单纯形消息传递提升语言模型困惑度的研究。公开的GPT-2 Large(6.2倍参数)零样本困惑度为22.05,作为外部参考,非本研究基准。三个负结果(层光滑、伴随往返、曲率正则化)及结构先验与精度加权联合结果共同支持“结构/一致性区分”经验模式:增加新拓扑的范畴先验有益,而强制一致性身份的先验无效。

原文摘要 · Abstract (English)

The Cognitive Categorical Transformer (CCT) is a 306M-parameter architecture that augments a pretrained GPT-2 Small backbone with cognitively grounded components derived from category theory and several inspirations from cognitive science. Under a matched-step protocol (215,000 optimizer steps, matched data, matched optimizer and schedule) on WikiText-103, CCT reaches 21.27 validation perplexity, compared with 24.19 for an identically fine-tuned GPT-2 Small baseline. The architecture therefore contributes a 2.92 PPL (12% relative) reduction beyond what in-domain fine-tuning alone provides. A retrain-from-scratch ablation that holds GT-Full simplicial message passing bypassed across the entire seven-phase activation schedule reaches 23.72 PPL, localizing 84% of the architectural improvement (2.45 of 2.92 PPL) to GT-Full. We present the first ablation-validated evidence that simplicial message passing improves language-model perplexity at the 306M-parameter scale on WikiText-103. Published GPT-2 Large reaches 22.05 zero-shot PPL on WikiText-103 with 6.2x more parameters than GPT-2 Small; this paper treats that number as an external published reference, not as the architectural benchmark. Three negative results on consistency-style categorical priors (sheaf smoothing, adjunction round-trip, curvature regularization) and the joint structural-prior result for GT-Full and PrecisionWeightedPP together support an empirical pattern termed the *structure/consistency distinction*, in which categorical priors that add new topology improve language modeling and those that enforce a consistency identity do not.

语言模型范畴论结构先验消融实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。