arXiv:2604.21265cs.CL2026-04

先学音乐再学语言,能显著提升小模型的语言学习效率。

Listen and Chant Before You Read: The Ladder of Beauty in LM Pre-Training

  • 用钢琴演奏数据预训练,再逐步过渡到诗歌、散文,提升语言建模能力。
  • 相比随机初始化,困惑度降低17.5%,且在不同模型规模下持续保持优势。
  • 适合研究小模型高效预训练或跨模态迁移学习的学者参考。

我们发现,在语言之前用音乐对Transformer进行预训练,能显著加速语言习得。采用钢琴演奏数据集MAESTRO,通过音乐→诗歌→散文的渐进式训练流程,相比随机初始化实现17.5%的困惑度下降(p < 0.001,5个种子),其中音乐和诗歌分别优化模型的内部计算与嵌入层。收敛测试表明该优势非短暂:在模型维度d=64时,多种子验证显示在收敛后仍保持5.5%的差距(p = 0.017),且每轮均更快收敛至更低损失。真实音乐表现与合成模式相当,仅需三分之一数据即可达到上限;缩放实验揭示最优预训练数据量随模型容量变化:从d=16到d=64,大数据集带来的优势由-3%增至+6%。在所研究范围(d ∈ {16,32,64},参数量达~400K)内,结果表明数据选择应依模型容量调整,且人类创作的结构化输出可作为小语言模型的高效预训练基础;对现代预训练规模下的更强结论,仍需更大规模实验支持。

原文摘要 · Abstract (English)

We show that pre-training a Transformer on music before language significantly accelerates language acquisition. Using piano performances (MAESTRO dataset), a developmental pipeline -- music $\to$ poetry $\to$ prose -- yields a $17.5\%$ perplexity improvement over random initialization ($p < 0.001$, 5 seeds), with music and poetry improving orthogonal model components (internal computation and embeddings, respectively). Convergence tests confirm that this is not a transient head start: at $d\!=\!64$, multi-seed validation (5 seeds) shows a persistent 5.5\% gap at plateau ($p = 0.017$), with the pipeline converging faster and to a lower loss in every run. Real music matches the transfer ceiling of synthetic patterns with one-third the data, and scaling experiments reveal that optimal pre-training data volume shifts with model capacity ($-3\% \to +3\% \to +6\%$ advantage of larger datasets from $d\!=\!16$ to $d\!=\!64$). Across the scales we study ($d\!\in\!\{16,32,64\}$, up to ${\sim}400$K parameters), these results suggest a capacity-dependent data curation principle and indicate that structured human creative outputs can provide an efficient pre-training substrate for small language models; stronger conclusions at modern pre-training scale will require substantially larger experiments.

预训练小模型音乐迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。