首个开源克什米尔语语音合成系统,显著提升低资源语言语音质量
Bolbosh: Script-Aware Flow Matching for Kashmiri Text-to-Speech
- 基于最优传输流匹配的监督跨语言迁移,解决配对数据少问题
- 引入去混响/静音裁剪/音量归一化三阶段增强,统一异构语音数据
- 显式编码克什米尔字母表,保留元音细微差异,适合方言敏感语言
克什米尔语由约700万人使用,但语音技术严重滞后,尽管其为官方语言且有丰富语言遗产。现有语音合成系统难以满足数字可访问性需求。本文提出首个专用于克什米尔语的开源神经语音合成系统Bolbosh。零样本多语言基线在该语言上仅得1.86分(MOS),主要因无法准确建模波斯-阿拉伯文变音符号和语言特有音系结构。为此,我们采用基于最优传输条件流匹配(OT-CFM)的监督跨语言适应策略,在Matcha-TTS框架内实现有限配对数据下的稳定对齐。进一步设计三阶段声学增强流程:去混响、静音裁剪、音量归一化,以统一异构语音源并稳定对齐学习。模型词汇表扩展为显式编码克什米尔图稿字符,保留细粒度元音区分。最终系统取得3.63分(MOS)和3.73(MCD),显著优于多语言基线,确立克什米尔语语音合成新基准。结果表明,对书写系统敏感且有监督的流模型适配,对变音符号敏感的低资源语言语音合成至关重要。代码与数据已公开:https://github.com/gaash-lab/Bolbosh。
原文摘要 · Abstract (English)
Kashmiri is spoken by around 7 million people but remains critically underserved in speech technology, despite its official status and rich linguistic heritage. The lack of robust Text-to-Speech (TTS) systems limits digital accessibility and inclusive human-computer interaction for native speakers. In this work, we present the first dedicated open-source neural TTS system designed for Kashmiri. We show that zero-shot multilingual baselines trained for Indic languages fail to produce intelligible speech, achieving a Mean Opinion Score (MOS) of only 1.86, largely due to inadequate modeling of Perso-Arabic diacritics and language-specific phonotactics. To address these limitations, we propose Bolbosh, a supervised cross-lingual adaptation strategy based on Optimal Transport Conditional Flow Matching (OT-CFM) within the Matcha-TTS framework. This enables stable alignment under limited paired data. We further introduce a three-stage acoustic enhancement pipeline consisting of dereverberation, silence trimming, and loudness normalization to unify heterogeneous speech sources and stabilize alignment learning. The model vocabulary is expanded to explicitly encode Kashmiri graphemes, preserving fine-grained vowel distinctions. Our system achieves a MOS of 3.63 and a Mel-Cepstral Distortion (MCD) of 3.73, substantially outperforming multilingual baselines and establishing a new benchmark for Kashmiri speech synthesis. Our results demonstrate that script-aware and supervised flow-based adaptation are critical for low-resource TTS in diacritic-sensitive languages. Code and data are available at: https://github.com/gaash-lab/Bolbosh.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。