arXiv:2602.15675cs.CL2026-02

用大模型生成埃及阿拉伯语语音数据,解决其语音合成资源匮乏问题。

LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models

  • 用大模型生成埃及阿拉伯语文本,再合成语音并自动转写验证。
  • 构建38小时高质量埃及阿拉伯语语音数据集,支持多领域语音合成。
  • 开源数据集与微调模型,推动方言语音合成研究进展。

尽管神经文本到语音(TTS)技术取得进展,许多阿拉伯语方言仍缺乏资源支持,多数资料集中于现代标准阿拉伯语(MSA)和海湾方言,而最广泛理解的埃及阿拉伯语严重缺资源。本文提出NileTTS:包含38小时来自两名说话者、涵盖医疗、销售和通用对话等多样领域的转录语音数据。该数据集通过创新的合成管道构建:大语言模型(LLM)生成埃及阿拉伯语内容,经语音合成工具转换为自然语音,再通过自动转写与说话人分离,并辅以人工质量验证。我们基于XTTS v2——一种先进的多语言TTS模型,在该数据集上进行微调,并与在其他阿拉伯方言上训练的基线模型对比评估。贡献包括:(1) 首个公开可用的埃及阿拉伯语TTS数据集,(2) 可复现的方言语音合成数据生成流程,(3) 开源微调后的模型。所有资源均已发布,以推动埃及阿拉伯语语音合成研究。

原文摘要 · Abstract (English)

Despite the advances in neural text to speech (TTS), many Arabic dialectal varieties remain marginally addressed, with most resources concentrated on Modern Spoken Arabic (MSA) and Gulf dialects, leaving Egyptian Arabic -- the most widely understood Arabic dialect -- severely under-resourced. We address this gap by introducing NileTTS: 38 hours of transcribed speech from two speakers across diverse domains including medical, sales, and general conversations. We construct this dataset using a novel synthetic pipeline: large language models (LLM) generate Egyptian Arabic content, which is then converted to natural speech using audio synthesis tools, followed by automatic transcription and speaker diarization with manual quality verification. We fine-tune XTTS v2, a state-of-the-art multilingual TTS model, on our dataset and evaluate against the baseline model trained on other Arabic dialects. Our contributions include: (1) the first publicly available Egyptian Arabic TTS dataset, (2) a reproducible synthetic data generation pipeline for dialectal TTS, and (3) an open-source fine-tuned model. All resources are released to advance Egyptian Arabic speech synthesis research.

语音合成大模型埃及阿拉伯语数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。