arXiv:2507.22612cs.SDcs.AI2025-07

提出自适应时长模型,提升文本到语音对齐精度与零样本语音合成鲁棒性。

Adaptive Duration Model for Text Speech Alignment

  • 设计可自适应条件的时长预测框架,生成细粒度音素级时长分布
  • 在音素对齐准确率上显著优于基线模型,零样本合成更抗音频不匹配
  • 适用于需要高精度语音对齐的端到端语音合成任务

语音到文本对齐是神经文本到语音(TTS)模型的关键组件。自回归TTS模型通常使用注意力机制在线学习对齐,而非自回归端到端TTS模型则依赖外部来源提取的时长。本文提出一种新颖的时长预测框架,可根据给定文本生成具有前景的音素级时长分布。实验表明,所提模型在预测精度和条件自适应能力上均优于先前基线模型。具体而言,其在音素级对齐准确率上取得显著提升,并增强了零样本TTS模型在提示音频与输入音频不匹配情况下的性能鲁棒性。

原文摘要 · Abstract (English)

Speech-to-text alignment is a critical component of neural text to speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-line, while non-autoregressive end to end TTS models rely on durations extracted from external sources. In this paper, we propose a novel duration prediction framework that can give promising phoneme-level duration distribution with given text. In our experiments, the proposed duration model has more precise prediction and adaptation ability to conditions, compared to previous baseline models. Specifically, it makes a considerable improvement on phoneme-level alignment accuracy and makes the performance of zero-shot TTS models more robust to the mismatch between prompt audio and input audio.

语音合成时长预测对齐优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。