arXiv:2510.13915cs.CLcs.AI2025-10中稿 · COLM

小模型学得快,不靠语言简单,而靠文本统计规律。

Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models

  • 用结构相似但难易不同的合成数据测试小模型学习能力
  • 复杂文本训练的模型比简单文本更快产生连贯输出
  • n-gram多样性比可读性更能预测小模型能否学会

近期研究发现,极小的语言模型(SLMs)在儿童读物风格的数据集(如TinyStories)上训练时,能生成相当连贯的文本。这被解释为可读性(词汇简单、句式清晰、叙事熟悉)促进了小模型能力涌现。本文挑战这一观点:我们构建了结构一致但可读性不同的合成数据集,发现可读性本身并不能预测模型的连贯性或学习效率。在成人级复杂文本上训练的模型表现与简化文本相当,甚至在训练初期更快出现连贯输出。相反,我们发现统计上的简单性——以n-gram多样性衡量——是更强的学习能力预测因子。研究警示不应盲目类比人类认知发展,强调需基于实证推断小模型能力涌现的关键因素。

原文摘要 · Abstract (English)

Recent studies suggest that very small language models (SLMs) can generate surprisingly coherent text when trained on simplified, child-directed corpora such as TinyStories. These findings have been interpreted as evidence that readability -- characterized by accessible vocabulary, familiar narrative structure, and simple syntax -- plays a key role in enabling such capabilities to emerge. In this paper, we challenge that interpretation. We construct synthetic datasets with matched structure but varied readability, and find that readability alone does not predict coherence or learning efficiency in SLMs. Models trained on complex, adult-level text perform comparably to those trained on simplified language, and even exhibit faster development of coherence during training. Instead, we show that statistical simplicity, as measured by n-gram diversity, is a stronger predictor of learnability. Our findings caution against the growing trend of anthropomorphizing language model training -- drawing parallels to human cognitive development without empirical basis -- and argue for more precise reasoning about what properties actually support capability emergence in small models.

小模型语言模型可读性学习效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。