arXiv:2604.02043cs.CLcs.AI2026-04被引 1

揭示自监督语音模型学习语言结构的演化过程

Tracking the emergence of linguistic structure in self-supervised models learning from speech

  • 分析六种模型在不同训练阶段的层级结构编码规律
  • 发现语言结构在各层中呈现差异化的学习轨迹
  • 适合研究语音表征与语言学关系的学者参考

自监督语音模型能有效捕捉口语的语言结构,但这些结构在训练过程中何时出现?我们研究了六种在荷兰语语音数据上训练的Wav2Vec2和HuBERT模型,在不同层及中间检查点中对多种语言结构的编码情况。结果表明,不同层次的语言结构表现出显著不同的层间模式和学习轨迹,这可部分归因于它们从声学信号中抽象的程度以及输入信息整合的时间尺度差异。此外,预训练目标定义的层级显著影响语言结构的层间组织和学习路径,更高阶的预测任务(如迭代优化的伪标签)会引发更强的并行性。

原文摘要 · Abstract (English)

Self-supervised speech models learn effective representations of spoken language, which have been shown to reflect various aspects of linguistic structure. But when does such structure emerge in model training? We study the encoding of a wide range of linguistic structures, across layers and intermediate checkpoints of six Wav2Vec2 and HuBERT models trained on spoken Dutch. We find that different levels of linguistic structure show notably distinct layerwise patterns as well as learning trajectories, which can partially be explained by differences in their degree of abstraction from the acoustic signal and the timescale at which information from the input is integrated. Moreover, we find that the level at which pre-training objectives are defined strongly affects both the layerwise organization and the learning trajectories of linguistic structures, with greater parallelism induced by higher-order prediction tasks (i.e. iteratively refined pseudo-labels).

自监督学习语音表征语言结构模型演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。