通过音色克隆对比,自动发现发音错误。
Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison
- 用语音克隆生成标准发音版本,逐帧比对差异。
- 无需预设音素规则,可在少数据下检测错误。
- 适合语言学习者或语音矫正工具开发者使用。
本文提出一种新方法,通过分析用户原始语音与经过修正发音的语音克隆版本之间的声学差异,来检测发音错误。我们假设原始语音与克隆语音之间声学差异最大的区域,可能对应于发音错误段落。该方法利用最新的语音克隆技术,生成具有用户音色但发音正确的合成语音,并进行逐帧比较,从而定位问题语音片段。实验表明,该方法能有效识别具体发音错误,且无需依赖预先定义的音素规则,也不需要为每种目标语言准备大量训练数据。
原文摘要 · Abstract (English)
This paper presents a novel approach for detecting mispronunciations by analyzing deviations between a user's original speech and their voice-cloned counterpart with corrected pronunciation. We hypothesize that regions with maximal acoustic deviation between the original and cloned utterances indicate potential mispronunciations. Our method leverages recent advances in voice cloning to generate a synthetic version of the user's voice with proper pronunciation, then performs frame-by-frame comparisons to identify problematic segments. Experimental results demonstrate the effectiveness of this approach in pinpointing specific pronunciation errors without requiring predefined phonetic rules or extensive training data for each target language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。