构建并评估了用于电信客服的合成孟加拉语语音数据集。
Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care
- 用OmniVoice生成1万条语音文本对,支持语音克隆与高精度合成。
- 自动评估显示字错误率2.54%,词错误率0.59%,一致性良好。
- 适合语音识别、合成模型训练及低资源语言研究者使用。
面向客户交互应用的语音系统通常需要特定领域的语言覆盖。本文构建了一个用于电信客服场景的合成孟加拉语语音数据集,包含10,000个音视频对,约26.82小时的24 kHz语音,并按9,000:500:500划分训练、验证和测试集。该数据集在Hugging Face上以CC-BY-4.0许可证公开。语音采用OmniVoice在语音克隆模式下生成,基于真实女性参考录音与文本,使用bfloat16精度、16步扩散采样和说话速率控制值1.0。除原始孟加拉语文本外,还提供经标准化处理的转录字段,适用于语音识别/语音转写(ASR/STT)训练与评估。我们利用微调自bengaliAI/tugstugi_bengaliai-regional-asr_whisper-medium的领域适配型Whisper ASR模型对全部10,000样本进行自动可懂度检查,并对部分样本进行人工听觉评估。结果表明平均词错误率(WER)为2.54%,平均字符错误率(CER)为0.59%,中位数WER与CER均为0.00%。这些结果表明在所选自动评估流程下,文本与语音间具有强一致性。同时论文也讨论了合成语音及基于语音转写评估的局限性。
原文摘要 · Abstract (English)
Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs, approximately 26.82 hours of 24 kHz speech, and predefined train, validation, and test splits of 9,000, 500, and 500 examples. It is publicly released on Hugging Face under the CC-BY-4.0 license. The speech was generated with OmniVoice in voice-cloning mode using a real female reference recording and transcript, with bfloat16 precision, 16 diffusion sampling steps, and a speaking-rate control value of 1.0. Along with the original Bengali text, the dataset provides a normalized transcript field designed for ASR/STT training and evaluation. We report an automatic intelligibility check over all 10,000 samples using a domain-adapted Whisper ASR model fine-tuned from bengaliAI/tugstugi_bengaliai-regional-asr_whisper-medium, along with a manual listening check on selected samples. The evaluation gives an average WER of 2.54%, an average CER of 0.59%, and median WER and CER values of 0.00%. These results suggest strong text-audio consistency under the selected automatic evaluation pipeline, while the paper also discusses the limitations of synthetic speech and STT-based evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。