arXiv:2510.10774cs.SDcs.AI2025-10被引 2

构建了首个大规模波斯语多说话人语音数据集,推动低资源语音合成发展。

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis

  • 基于播客录音开发自动化数据构建流水线,融合语言模型与语音识别优化对齐
  • 产出2200小时高质量数据,含136万段语音文本对,覆盖1815个说话人
  • 适合作为波斯语语音合成研究的基础数据集,尤其适合零样本多语言模型训练

波斯语在公开语音-文本资源中仍严重缺失,制约了多说话人语音合成、语音语言建模及低资源语音处理的发展。本文提出ParsVoice,目前最大且公开可用的波斯语语音-文本语料库,专为训练多说话人语音合成系统设计,并提供可扩展的数据构建流程,从长篇有声书录音中生成高质量语音-文本数据。该流程结合微调的ParsBERT句子补全分类器、基于ASR的边界优化、标点恢复、说话人识别以及涵盖音频与波斯语特有文本属性的多维度质量评估。最终释放的2200小时语音-文本子集包含136万条对齐片段,来自1815个自动识别的说话人身份,规模超过此前最大开源波斯语语音合成数据集25倍以上。为验证数据质量,我们微调了零样本多语言语音合成模型XTTS(直接处理原始波斯文,无需音素表示),达到自然度3.6/5分,说话人相似度4.0/5分。ParsVoice数据集已公开:https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice。

原文摘要 · Abstract (English)

Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech processing. We introduce ParsVoice, the largest publicly available Persian speech-text corpus tailored for training multi-speaker TTS systems, along with a scalable pipeline to construct high-quality speech-text data from long-form audiobook recordings. The pipeline combines a fine-tuned ParsBERT sentence-completion classifier, ASR-based boundary optimization, punctuation restoration, speaker identification, and a multi-dimensional quality assessment that covers both audio and Persian-specific text properties. The resulting release contains a 2,200-hour TTS-ready subset with 1.36 million aligned segments from 1,815 automatically identified speaker IDs, making it more than 25 times larger than the previously largest open Persian TTS dataset. To validate the corpus, we fine-tune XTTS, a zero-shot multilingual TTS model that operates directly on raw Persian text without phoneme representations, achieving a naturalness MOS of 3.6/5 and speaker similarity MOS of 4.0/5. The ParsVoice dataset is publicly available at: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice.

语音合成多说话人波斯语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。