针对满语数据少、构词复杂难题,提出分层文本表示与流匹配合成新方法。
ManchuTTS: Towards High-Quality Manchu Speech Synthesis via Flow Matching and Hierarchical Text Representation
- 设计三级文本表征与跨模态分层注意力,对齐音素、音节与语调
- 在5.2小时数据上实现4.52的听感评分,比基线模型显著更优
- 适合濒危语言语音合成研究者,尤其关注构词复杂语言
作为濒危语言,满语在语音合成中面临数据稀缺和强形态黏着性两大挑战。本文提出面向满语的语音合成系统ManchuTTS,针对其语言特征设计三层文本表征(音素、音节、韵律)及跨模态分层注意力机制,实现多粒度对齐。合成模型结合深度卷积网络与流匹配Transformer,支持高效非自回归生成,并引入分层对比损失以引导声学-语言结构对应关系。为应对低资源问题,构建首个满语语音合成数据集,并采用数据增强策略。实验表明,使用由6.24小时标注语料中抽取的5.2小时子集训练,模型达到4.52的平均主观评分(MOS),显著优于所有基线模型。消融实验验证,分层指导使黏着词发音准确率(AWPA)提升31%,韵律自然度提升27%。
原文摘要 · Abstract (English)
As an endangered language, Manchu presents unique challenges for speech synthesis, including severe data scarcity and strong phonological agglutination. This paper proposes ManchuTTS(Manchu Text to Speech), a novel approach tailored to Manchu's linguistic characteristics. To handle agglutination, this method designs a three-tier text representation (phoneme, syllable, prosodic) and a cross-modal hierarchical attention mechanism for multi-granular alignment. The synthesis model integrates deep convolutional networks with a flow-matching Transformer, enabling efficient, non-autoregressive generation. This method further introduce a hierarchical contrastive loss to guide structured acoustic-linguistic correspondence. To address low-resource constraints, This method construct the first Manchu TTS dataset and employ a data augmentation strategy. Experiments demonstrate that ManchuTTS attains a MOS of 4.52 using a 5.2-hour training subset derived from our full 6.24-hour annotated corpus, outperforming all baseline models by a notable margin. Ablations confirm hierarchical guidance improves agglutinative word pronunciation accuracy (AWPA) by 31% and prosodic naturalness by 27%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。