用轻量微调让大模型同时实现发音评分与错误诊断
English Pronunciation Evaluation without Complex Joint Training: LoRA Fine-tuned Speech Multimodal LLM
- 仅微调LoRA层,无需复杂结构改动
- 评分相关性超0.7,词错率和音素错率均低于0.15
- 适合英语二语学习者使用的高效语音评估工具
本研究证明,通过低秩适配(LoRA)微调的多模态大模型(MLLM)可同时完成自动发音评估(APA)与发音错误检测诊断(MDD)。基于微软Phi-4-multimodal-instruct模型,该方法无需复杂的架构调整或独立训练流程。在Speechocean762数据集上微调后,模型预测的发音评分与人工评分的皮尔逊相关系数(PCC > 0.7),词错误率(WER)和音素错误率(PER)均低于0.15。值得注意的是,仅微调LoRA层即可达到与全音频层微调相当的性能。研究表明,无需全量微调即可构建集成式发音评估系统,相比以往联合模型更具效率,为英语二语学习者的计算机辅助发音训练(CAPT)技术提供了更便捷、整合且高效的解决方案。
原文摘要 · Abstract (English)
This study demonstrates that a Multimodal Large Language Model (MLLM) adapted via Low-Rank Adaptation (LoRA) can perform both Automatic Pronunciation Assessment (APA) and Mispronunciation Detection and Diagnosis (MDD) simultaneously. Leveraging Microsoft's Phi-4-multimodal-instruct, our fine-tuning method eliminates the need for complex architectural changes or separate training procedures conventionally required for these distinct tasks. Fine-tuned on the Speechocean762 dataset, the pronunciation evaluation scores predicted by the model exhibited a strong Pearson Correlation Coefficient (PCC > 0.7) with human-assigned scores, while achieving low Word Error Rate (WER) and Phoneme Error Rate (PER) (both < 0.15). Notably, fine-tuning only the LoRA layers was sufficient to achieve performance levels comparable to those achieved by fine-tuning all audio layers. This research highlights that an integrated pronunciation assessment system can be established by adapting large multimodal models without full fine-tuning, utilizing a significantly simpler training methodology compared to previous joint models designed for simultaneous APA and MDD. This efficient LoRA-based approach paves the way for more accessible, integrated, and effective Computer-Assisted Pronunciation Training (CAPT) technologies for English L2 learners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。