用语音转换模拟母语者跟读,帮二语学习者发现发音问题。
A Pilot Study of Applying Sequence-to-Sequence Voice Conversion to Evaluate the Intelligibility of L2 Speech Using a Native Speaker's Shadowings
- 用语音转换技术模拟母语者跟读非母语语音。
- 虚拟跟读输出在语言和声学上与真实跟读相似度较高。
- 适合语言教学系统开发与发音纠错研究。
二语说话人因发音错误和语调不当可能导致话语难以理解。当前计算机辅助语言学习系统通常依赖语音识别引擎提供文本反馈,但理想反馈应足够细致,帮助学习者识别并诊断发音问题。受语言教师通过声音对声音方式纠正发音的启发,本初步研究利用一个独特的半平行数据集,包含二语学习者朗读、母语者跟读及对应脚本跟读的语音。探索使用语音转换技术复制母语者跟读二语语音的过程,构建虚拟跟读系统。实验结果表明,该语音转换系统在模拟母语者跟读行为方面具有可行性,虚拟跟读输出在语言和声学层面均与真实母语跟读表现出合理相似性。
原文摘要 · Abstract (English)
Utterances by L2 speakers can be unintelligible due to mispronunciation and improper prosody. In computer-aided language learning systems, textual feedback is often provided using a speech recognition engine. However, an ideal form of feedback for L2 speakers should be so fine-grained that it enables them to detect and diagnose unintelligible parts of L2 speakers' utterances. Inspired by language teachers who correct students' pronunciation through a voice-to-voice process, this pilot study utilizes a unique semi-parallel dataset composed of non-native speakers' (L2) reading aloud, shadowing of native speakers (L1) and their script-shadowing utterances. We explore the technical possibility of replicating the process of an L1 speaker's shadowing L2 speech using Voice Conversion techniques, to create a virtual shadower system. Experimental results demonstrate the feasibility of the VC system in simulating L1's shadowing behavior. The output of the virtual shadower system shows a reasonable similarity to the real L1 shadowing utterances in both linguistic and acoustic aspects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。