研究时长预测对印度语语音合成的影响,发现不同策略各有优劣。
Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages
- 采用非自回归连续流模型,测试多种时长预测方法。
- 零样本生成中,部分语言用时长补全提升可懂度,另一些则靠说话人提示保真度。
- 适合需要高可懂或保真的低资源语言语音合成任务。
低资源语言(如许多印度语言)的高质量语音生成仍面临挑战,主要受限于数据稀缺和语言结构多样性。时长预测是语音生成流程中的关键组件,对韵律和语流建模至关重要。尽管一些近期生成方法选择省略显式时长建模,但常以更长训练时间为代价。本文保留并探索该模块,旨在深入理解其在语言丰富但数据匮乏的印度语环境下的作用。我们基于公开的印度语言数据,训练了一种基于非自回归连续归一化流(CNF)的语音模型,并评估了多种时长预测策略在零样本、说话人特定生成任务中的表现。在语音补全任务上的对比分析揭示出细微权衡:基于语音补全的预测器在某些语言中提升了可懂度,而基于说话人提示的预测器则在其他语言中更好地保留了说话人特征。这些发现为针对特定语言与任务设计时长策略提供了依据,强调了在低资源多语言场景下,时长预测等可解释组件在适配先进生成架构中的持续价值。
原文摘要 · Abstract (English)
High-quality speech generation for low-resource languages, such as many Indian languages, remains a significant challenge due to limited data and diverse linguistic structures. Duration prediction is a critical component in many speech generation pipelines, playing a key role in modeling prosody and speech rhythm. While some recent generative approaches choose to omit explicit duration modeling, often at the cost of longer training times. We retain and explore this module to better understand its impact in the linguistically rich and data-scarce landscape of India. We train a non-autoregressive Continuous Normalizing Flow (CNF) based speech model using publicly available Indian language data and evaluate multiple duration prediction strategies for zero-shot, speaker-specific generation. Our comparative analysis on speech-infilling tasks reveals nuanced trade-offs: infilling based predictors improve intelligibility in some languages, while speaker-prompted predictors better preserve speaker characteristics in others. These findings inform the design and selection of duration strategies tailored to specific languages and tasks, underscoring the continued value of interpretable components like duration prediction in adapting advanced generative architectures to low-resource, multilingual settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。