用对话式语音模型提升英语发音训练反馈质量
Unlocking Large Audio-Language Models for Interactive Language Learning
- 构建带详细纠错与改进建议的英语发音数据集
- 指令微调后模型在纠错与建议生成上显著超越基线
- 适合语言学习系统开发者与教育AI研究者
第二语言发音训练仍面临挑战,尽管已有计算机辅助发音训练(CAPT)系统。传统CAPT系统提供的反馈往往不直观且缺乏可操作性,限制了其效果。音频-语言模型(ALMs)的发展为改进此类系统提供了新可能。本文通过引入L2-Arctic-plus数据集,该数据集包含详细的发音错误解释和可操作的改进建议,探索基于对话的发音训练中ALMs的应用。我们在该数据集上对级联ASR+LLM与现有ALMs进行基准测试,重点评估其在误发音检测和生成可操作反馈方面的能力。为进一步提升性能,我们提出在L2-Arctic-plus上对ALMs进行指令微调。实验结果表明,指令微调后的模型在客观指标和人工评估上均显著优于现有基线,验证了所提数据集的价值。
原文摘要 · Abstract (English)
Achieving pronunciation proficiency in a second language (L2) remains a challenge, despite the development of Computer-Assisted Pronunciation Training (CAPT) systems. Traditional CAPT systems often provide unintuitive feedback that lacks actionable guidance, limiting its effectiveness. Recent advancements in audio-language models (ALMs) offer the potential to enhance these systems by providing more user-friendly feedback. In this work, we investigate ALMs for chat-based pronunciation training by introducing L2-Arctic-plus, an English dataset with detailed error explanations and actionable suggestions for improvement. We benchmark cascaded ASR+LLMs and existing ALMs on this dataset, specifically in detecting mispronunciation and generating actionable feedback. To improve the performance, we further propose to instruction-tune ALMs on L2-Arctic-plus. Experimental results demonstrate that our instruction-tuned models significantly outperform existing baselines on mispronunciation detection and suggestion generation in terms of both objective and human evaluation, highlighting the value of the proposed dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。