arXiv:2410.22906cs.CL2024-10被引 13

用音素流预训练语言模型,探索语音层面的语义理解新路径

From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes

  • 将文本数据转换为连续音素流,实现音素级语言模型预训练
  • 在传统任务上性能略有下降,但提升语音相关任务的可解释性
  • 适合研究语音认知、语音生成与音素建模的学者与工程师

语言模型通常以文本形式进行大规模预训练,但将数据表示为音素流也具有独特优势,如深入理解语音习得机制或提升语音相关任务表现。挑战在于缺乏对应的音素评估基准。为此,我们开发了一套将文本数据转化为连续音素流的管道,应用于婴儿语言模型挑战赛中的1亿词预训练数据集以及标准语言与语法测试集,实现了基于音素输入的模型预训练与评估。结果表明,音素训练在传统语言理解任务上性能略有下降,但在语音分析和实际应用中展现出显著价值。

原文摘要 · Abstract (English)

Language models are typically trained on large corpora of text in their default orthographic form. However, this is not the only option; representing data as streams of phonemes can offer unique advantages, from deeper insights into phonological language acquisition to improved performance on sound-based tasks. The challenge lies in evaluating the impact of phoneme-based training, as most benchmarks are also orthographic. To address this, we develop a pipeline to convert text datasets into a continuous stream of phonemes. We apply this pipeline to the 100-million-word pre-training dataset from the BabyLM challenge, as well as to standard language and grammatical benchmarks, enabling us to pre-train and evaluate a model using phonemic input representations. Our results show that while phoneme-based training slightly reduces performance on traditional language understanding tasks, it offers valuable analytical and practical benefits.

音素建模语言模型语音理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。