开源中英文双轨对话数据集,提升语音合成自然度。
Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis
- 构建中英双语15小时真实对话数据,分离声道保质
- 微调后语音合成在主观与客观评测中全面优于基线
- 适合语音合成、人机交互研究者使用
全双工自发对话数据对提升对话式语音合成的自然性与交互性至关重要。本文发布两个开源双轨对话语音数据集,分别涵盖中文和英文,总计15小时,均在隔离房间录制,每名说话人拥有独立高质量音频轨道。数据覆盖多样化日常话题与场景,包含频繁重叠、回应词、笑声等真实互动行为。文中详述数据采集流程、转写与标注方法。通过在基线语音合成模型上微调,新模型在主观与客观评估中均表现更优,证明合成语音自然度与对话真实感显著提升。所有数据、标注及微调评估代码均已公开,推动对话语音合成研究发展。
原文摘要 · Abstract (English)
Full-duplex, spontaneous conversational data are essential for enhancing the naturalness and interactivity of synthesized speech in conversational TTS systems. We present two open-source dual-track conversational speech datasets, one in Chinese and one in English, designed to enhance the naturalness of synthesized speech by providing more realistic conversational data. The two datasets contain a total of 15 hours of natural, spontaneous conversations recorded in isolated rooms, which produces separate high-quality audio tracks for each speaker. The conversations cover diverse daily topics and domains, capturing realistic interaction patterns including frequent overlaps, backchannel responses, laughter, and other non-verbal vocalizations. We introduce the data collection procedure, transcription and annotation methods. We demonstrate the utility of these corpora by fine-tuning a baseline TTS model with the proposed datasets. The fine-tuned TTS model achieves higher subjective and objective evaluation metrics compared to the baseline, indicating improved naturalness and conversational realism in synthetic speech. All data, annotations, and supporting code for fine-tuning and evaluation are made available to facilitate further research in conversational speech synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。