arXiv:2601.13802cs.CLcs.SD2026-01被引 2

首个统一阿拉伯多方言语音合成框架,解决数据少与评测难问题。

Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis

  • 用多步骤清洗流程将开源语音识别数据转为多方言语音训练集。
  • 零样本合成效果超越单方言模型,支持无分词文本直接生成。
  • 开源完整代码、模型和首个多方言评测基准,适合研究者使用。

阿拉伯语有超过30种口语变体,但尚无开源的统一语音合成系统。主要障碍包括方言间词汇与发音差异大、高质量合成数据稀缺,以及缺乏标准化的多方言评估基准。我们提出Habibi,一个统一阿拉伯多方言语音合成框架,解决了上述三大挑战。通过多阶段数据清洗流程,我们将开源语音识别语料库转化为覆盖12种以上地区方言的语音合成训练数据。采用语言学启发的渐进式课程学习策略——从现代标准阿拉伯语逐步过渡到方言数据——使模型无需文本分词即可实现鲁棒的零样本合成。我们还发布了首个标准化的多方言阿拉伯语语音合成评测基准,包含7个方言子集,共超过11,000条经人工验证的语音片段。在该基准上,统一模型表现达到或优于各方言专用模型。自动指标与人工评测均表明,Habibi在可懂度、说话人相似性和自然度方面与ElevenLabs Eleven v3(alpha)相当。大规模消融实验(约8,000小时H100 GPU计算,30+配置)验证了每个设计选择的有效性。我们已将所有检查点、训练与推理代码及基准数据开源,地址为https://SWivid.github.io/Habibi/,这是首个针对多方言阿拉伯语语音合成的完整开源发布。

原文摘要 · Abstract (English)

Arabic spans over 30 spoken varieties, yet no open-source text-to-speech system unifies them. Key barriers include substantial cross-dialect lexical and phonological divergence, scarce synthesis-grade data, and the absence of a standardized multi-dialect evaluation benchmark. We present Habibi, a unified-dialectal Arabic TTS framework that addresses all three. Through a multi-step curation pipeline, we repurpose open-source ASR corpora into TTS training data covering 12+ regional dialects. A linguistically-informed curriculum learning strategy - progressing from Modern Standard Arabic to dialectal data - enables robust zero-shot synthesis without text diacritization. We further release the first standardized multi-dialect Arabic TTS benchmark, comprising over 11,000 utterances across 7 dialect subsets with manually verified transcripts. On this benchmark, our unified model matches or surpasses per-dialect specialized models. Both automatic metrics and human evaluations confirm that Habibi is highly competitive with ElevenLabs' Eleven v3 (alpha) in intelligibility, speaker similarity, and naturalness. Extensive ablations (~8,000 H100 GPU hours, 30+ configurations) validate each design choice. We open-source all checkpoints, training and inference code, and benchmark data - the first such release for multi-dialect Arabic TTS - at https://SWivid.github.io/Habibi/ .

语音合成多方言开源阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。