用音标统一方言发音,零样本快速生成新方言语音。
DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation
- 基于国际音标构建统一发音框架,解决拼写不一致问题。
- 引入专家混合模型与低秩适配器,仅需数小时数据即实现零样本迁移。
- 适合方言语音合成、文化表达数字化等应用,无需大量标注数据。
方言语音承载丰富的文化和语言多样性,但因数据稀缺、拼写不统一及发音复杂,构建方言文语转换(TTS)系统仍具挑战。为此,我们提出DiaMoE-TTS,一种基于国际音标(IPA)的统一框架,标准化发音表示并解决音形转换歧义。该系统基于F5-TTS架构,引入方言感知的专家混合(MoE)模型以建模音系差异,并采用低秩适配器(LoRA)与条件适配器实现参数高效适应,可快速迁移到新方言。相比依赖大规模或专有资源的方法,DiaMoE-TTS支持可扩展的开源数据驱动合成。实验表明,其能生成自然流畅的语音,在未见方言及特定领域(如京戏)上仅需数小时数据即可实现零样本表现。
原文摘要 · Abstract (English)
Dialect speech embodies rich cultural and linguistic diversity, yet building text-to-speech (TTS) systems for dialects remains challenging due to scarce data, inconsistent orthographies, and complex phonetic variation. To address these issues, we present DiaMoE-TTS, a unified IPA-based framework that standardizes phonetic representations and resolves grapheme-to-phoneme ambiguities. Built upon the F5-TTS architecture, the system introduces a dialect-aware Mixture-of-Experts (MoE) to model phonological differences and employs parameter-efficient adaptation with Low-Rank Adaptors (LoRA) and Conditioning Adapters for rapid transfer to new dialects. Unlike approaches dependent on large-scale or proprietary resources, DiaMoE-TTS enables scalable, open-data-driven synthesis. Experiments demonstrate natural and expressive speech generation, achieving zero-shot performance on unseen dialects and specialized domains such as Peking Opera with only a few hours of data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。