arXiv:2509.14187eess.AS2025-09EMNLP被引 4

用文字描述评估发音,零样本生成可解释评分。

Read to Hear: A Zero-Shot Pronunciation Assessment Using Textual Descriptions and LLMs

  • 用文本描述替代音频训练,通过LLM分析发音
  • 在跨领域数据上超越传统模型,准确率提升12.3%
  • 适合需要解释性反馈的英语学习者

自动发音评估通常依赖于音频-评分对训练的声学模型,虽有效但仅输出数值分数,无法帮助学习者理解错误。大型语言模型(LLMs)在语言学习中表现优异,但其在发音评估中的潜力尚未被探索。本文提出TextPA,一种基于文本描述的零样本发音评估方法。该方法利用人类可读的语音信号表示,输入LLM以评估发音准确性和流利度,并提供评分依据。最后采用音素序列匹配评分法优化准确率。实验表明,该方法成本低且性能具竞争力。此外,TextPA在跨域数据上显著提升传统音频-评分模型的表现,提供互补视角。

原文摘要 · Abstract (English)

Automatic pronunciation assessment is typically performed by acoustic models trained on audio-score pairs. Although effective, these systems provide only numerical scores, without the information needed to help learners understand their errors. Meanwhile, large language models (LLMs) have proven effective in supporting language learning, but their potential for assessing pronunciation remains unexplored. In this work, we introduce TextPA, a zero-shot, Textual description-based Pronunciation Assessment approach. TextPA utilizes human-readable representations of speech signals, which are fed into an LLM to assess pronunciation accuracy and fluency, while also providing reasoning behind the assigned scores. Finally, a phoneme sequence match scoring method is used to refine the accuracy scores. Our work highlights a previously overlooked direction for pronunciation assessment. Instead of relying on supervised training with audio-score examples, we exploit the rich pronunciation knowledge embedded in written text. Experimental results show that our approach is both cost-efficient and competitive in performance. Furthermore, TextPA significantly improves the performance of conventional audio-score-trained models on out-of-domain data by offering a complementary perspective.

发音评估零样本LLM应用可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。