arXiv:2606.17662eess.AS2026-06

用合成语音提升印地语等三种语言的语音识别效果

An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages

  • 用真实与合成语音混合训练,提升印地语等三语识别性能
  • 合成语音中使用不同语音克隆数量影响识别准确率
  • 研究结果对低资源语言语音系统优化有指导意义

合成数据在机器学习模型训练中具有潜力,尤其在自动语音识别(ASR)系统中;然而其有效性需系统评估。本研究针对印地语、卡纳达语和泰卢固语三种印地语族语言,考察将合成语音数据与真实录音结合对ASR性能的影响。分析了真实数据与合成数据融合带来的性能提升,并独立研究了生成合成语音时脚本来源对性能的影响。同时评估了不同语音合成模型生成的合成语音效果。最后研究了语音克隆技术在合成语音生成中的影响,包括生成数据时使用不同数量克隆语音对识别性能的影响。

原文摘要 · Abstract (English)

Synthetic data has the potential to be a valuable resource for training machine learning models, particularly Automatic Speech Recognition (ASR) Systems; however, its effectiveness requires systematic evaluation. In this study, we investigate the impact of incorporating synthetic speech data alongside real-world recordings for three Indic languages: Hindi, Kannada, and Telugu. We analyze the performance gains achieved by augmenting synthetic data with real data and independently examine how ASR performance varies with the sources of scripts used to generate synthetic speech. In addition, we evaluate the effect of synthetic speech generated using different speech synthesis models. Finally, we study the impact of voice cloning in synthetic speech generation on ASR performance, including how performance varies with the number of distinct cloned voices used during data generation.

语音识别合成数据低资源语言语音克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。