arXiv:2601.16230eess.AScs.AI2026-01被引 3

语音大模型零样本评估二语发音,效果接近真人评分。

Zero-Shot Speech LLMs for Multi-Aspect Evaluation of L2 Speech: Challenges and Opportunities

  • 用指令微调的语音大模型直接评估发音四维度。
  • 对高质量语音评分与真人误差在±2以内。
  • 适合语音教学系统开发与自动评分研究者。

准确评估二语英语发音对语言学习至关重要,可提供个性化反馈并公平衡量进步。然而,由于句子级流利度、语调和完整性复杂,自动化评分仍具挑战性。本文评估了经过指令微调的语音大模型 Qwen2-Audio-7B-Instruct 在 5,000 条 Speechocean762 语句上的零样本表现。该模型生成符合评分量规的准确性、流利度、语调和完整性分数,与人类评分在 ±2 容差范围内具有强一致性,尤其在高质量语音上表现优异。但对低质量语音存在评分偏高问题,且错误检测精度不足。结果表明语音大模型在可扩展发音评估中具有巨大潜力,未来可通过优化提示工程、校准策略与语音特征融合进一步提升计算机辅助发音训练水平。

原文摘要 · Abstract (English)

An accurate assessment of L2 English pronunciation is crucial for language learning, as it provides personalized feedback and ensures a fair evaluation of individual progress. However, automated scoring remains challenging due to the complexity of sentence-level fluency, prosody, and completeness. This paper evaluates the zero-shot performance of Qwen2-Audio-7B-Instruct, an instruction-tuned speech-LLM, on 5,000 Speechocean762 utterances. The model generates rubric-aligned scores for accuracy, fluency, prosody, and completeness, showing strong agreement with human ratings within +-2 tolerance, especially for high-quality speech. However, it tends to overpredict low-quality speech scores and lacks precision in error detection. These findings demonstrate the strong potential of speech LLMs in scalable pronunciation assessment and suggest future improvements through enhanced prompting, calibration, and phonetic integration to advance Computer-Assisted Pronunciation Training.

语音大模型发音评估自动评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。