BnTTS让低资源孟加拉语语音合成仅用少量数据就能适配新发音人。
BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting
- 基于XTTS架构,针对孟加拉语音系特点改进多语言语音合成流程。
- 在3.85千小时数据上预训练,少样本设置下显著提升语音自然度与发音人保真度。
- 适合低资源语言语音合成研究者,尤其关注孟加拉语或跨语言迁移应用。
本文提出BnTTS(孟加拉语文本转语音),首个基于发音人适配的孟加拉语语音合成框架,旨在解决该语言在极低数据条件下的语音合成空白。在XTTS架构基础上,将孟加拉语融入多语言语音合成流程,并针对其音系与语言特征进行调整。模型在3.85千小时的孟加拉语语音数据集上进行预训练,使用自建测试集评估零样本与少样本场景下的性能。实证结果表明,在少样本设置下,BnTTS显著提升了合成语音的自然度、可懂度与发音人保真度。相比现有最优孟加拉语语音合成系统,其在主观平均意见分(SMOS)、自然度与清晰度指标上表现更优。
原文摘要 · Abstract (English)
This paper introduces BnTTS (Bangla Text-To-Speech), the first framework for Bangla speaker adaptation-based TTS, designed to bridge the gap in Bangla speech synthesis using minimal training data. Building upon the XTTS architecture, our approach integrates Bangla into a multilingual TTS pipeline, with modifications to account for the phonetic and linguistic characteristics of the language. We pre-train BnTTS on 3.85k hours of Bangla speech dataset with corresponding text labels and evaluate performance in both zero-shot and few-shot settings on our proposed test dataset. Empirical evaluations in few-shot settings show that BnTTS significantly improves the naturalness, intelligibility, and speaker fidelity of synthesized Bangla speech. Compared to state-of-the-art Bangla TTS systems, BnTTS exhibits superior performance in Subjective Mean Opinion Score (SMOS), Naturalness, and Clarity metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。