arXiv:2506.04152eess.AS2025-06中稿 · Interspeech 2025被引 16

构建36.7小时高保真英语语音数据集,支持高质量零样本语音合成。

HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset

  • 基于LibriVox音频书构建,含22.05kHz与44.1kHz双采样率数据
  • 36.7小时22.05kHz数据训练出高保真零样本TTS模型
  • 提供元数据和质量过滤工具,适配多场景研究需求

本文介绍HiFiTTS-2,一个面向高带宽语音合成的大规模语音数据集。数据源自LibriVox有声书,包含约36.7千小时的英文语音(用于22.05 kHz训练),以及31.7千小时(用于44.1 kHz训练)。我们提出了完整的数据处理流程,包括带宽估计、分段、文本预处理及多说话人检测。数据集附带由该流程生成的详细语句与有声书元数据,使研究人员可使用数据质量筛选器,根据具体需求定制数据集。实验表明,该数据处理流程与数据集能有效支持在高带宽条件下训练高质量的零样本文语转换(TTS)模型。

原文摘要 · Abstract (English)

This paper introduces HiFiTTS-2, a large-scale speech dataset designed for high-bandwidth speech synthesis. The dataset is derived from LibriVox audiobooks, and contains approximately 36.7k hours of English speech for 22.05 kHz training, and 31.7k hours for 44.1 kHz training. We present our data processing pipeline, including bandwidth estimation, segmentation, text preprocessing, and multi-speaker detection. The dataset is accompanied by detailed utterance and audiobook metadata generated by our pipeline, enabling researchers to apply data quality filters to adapt the dataset to various use cases. Experimental results demonstrate that our data pipeline and resulting dataset can facilitate the training of high-quality, zero-shot text-to-speech (TTS) models at high bandwidths.

语音合成高保真数据集TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。