arXiv:2507.20091cs.CLeess.AS2025-07被引 4

提出新语音分词方法,让语音大模型自动学会语气处理能力。

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

  • 用词级韵律标记替代传统语音分词,保留完整语调信息。
  • 仅靠预训练就实现语气对比、情绪识别和长文本韵律一致。
  • 适合研究语音生成、情感计算与语音理解的开发者参考。

语音语言模型具备语音处理与理解能力。一个关键期望是捕捉内容与韵律之间的复杂依赖关系。现有主流训练范式将语音转为离散标记后输入大语言模型,难以有效学习韵律信息——我们发现此类模型在仅通过预训练时,并未表现出明显的新兴韵律处理能力。为此,我们提出ProsodyLM,采用一种更利于韵律学习的简单分词方案:先将语音转写为文本,再添加一系列词级韵律标记。相比传统语音分词方式,该方案保留了更完整的韵律信息,且对基于文本的大语言模型更友好。我们发现,ProsodyLM仅通过预训练即可涌现出多样化的韵律处理能力,包括生成语音中的语气细微差别(如对比焦点)、理解语句中情绪与重音,以及在长上下文中保持韵律一致性。

原文摘要 · Abstract (English)

Speech language models refer to language models with speech processing and understanding capabilities. One key desirable capability for speech language models is the ability to capture the intricate interdependency between content and prosody. The existing mainstream paradigm of training speech language models, which converts speech into discrete tokens before feeding them into LLMs, is sub-optimal in learning prosody information -- we find that the resulting LLMs do not exhibit obvious emerging prosody processing capabilities via pre-training alone. To overcome this, we propose ProsodyLM, which introduces a simple tokenization scheme amenable to learning prosody. Each speech utterance is first transcribed into text, followed by a sequence of word-level prosody tokens. Compared with conventional speech tokenization schemes, the proposed tokenization scheme retains more complete prosody information, and is more understandable to text-based LLMs. We find that ProsodyLM can learn surprisingly diverse emerging prosody processing capabilities through pre-training alone, ranging from harnessing the prosody nuances in generated speech, such as contrastive focus, understanding emotion and stress in an utterance, to maintaining prosody consistency in long contexts.

语音生成韵律建模大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。