arXiv:2410.00025cs.CLcs.SD2024-10EMNLP被引 8

用音素分类微调语音模型,显著提升口语语言建模效果

Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach

  • 通过音素分类任务微调语音表征
  • 仅用少量数据达到文本模型百倍数据的效果
  • 适合追求高效语音建模的研究者

近年来,口语语言建模取得了进展,证明可直接从语音学习语言。传统文本级语音合成会丢失语调、语气等细节,而纯语音系统需比文本系统多出三个数量级的数据才能达到相似的语义能力。本文表明,在音素分类任务上微调语音表征模型,可获得更具上下文不变性的表示,基于这些单元训练的语言模型在词汇理解能力上可媲美使用百倍数据训练的文本模型。

原文摘要 · Abstract (English)

Recent progress in Spoken Language Modeling has shown that learning language directly from speech is feasible. Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and non-verbal vocalizations. Modeling directly from speech opens up the path to more natural and expressive systems. On the other hand, speech-only systems require up to three orders of magnitude more data to catch up to their text-based counterparts in terms of their semantic abilities. We show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data.

语音建模音素分类微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。