提出分层发音评估模型,可自动分析自由口语发音质量。
HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
- 构建多层级模型,从语音中自动评估学习者发音水平
- 在Speechocean762数据集上超越现有方法,评分更准确
- 适合需要实时反馈的口语教学场景
自动发音评估(APA)旨在通过提供及时、细粒度的诊断反馈,量化第二语言学习者在目标语言中的发音能力。现有研究主要聚焦于朗读类受限任务;而对自由口语场景下的发音质量评估仍相对不足。为此,本文提出HiPPO,一种面向口语语言的分层发音评估模型,仅基于学习者说出的语音,在多个语言层次上评估其口语水平。为提升评估准确性,引入对比序数正则化和课程学习策略进行模型训练:前者利用回归目标的序数特性生成可区分评分的特征,后者逐步增加训练复杂度,以适应自由口语输入的任务。在Speechocean762基准数据集上的实验验证了该方法的有效性与优越性,优于多种前沿基线模型。
原文摘要 · Abstract (English)
Automatic pronunciation assessment (APA) seeks to quantify a second language (L2) learner's pronunciation proficiency in a target language by offering timely and fine-grained diagnostic feedback. Most existing efforts on APA have predominantly concentrated on highly constrained reading-aloud tasks (where learners are prompted to read a reference text aloud); however, assessing pronunciation quality in unscripted speech (or free-speaking scenarios) remains relatively underexplored. In light of this, we first propose HiPPO, a hierarchical pronunciation assessment model tailored for spoken languages, which evaluates an L2 learner's oral proficiency at multiple linguistic levels based solely on the speech uttered by the learner. To improve the overall accuracy of assessment, a contrastive ordinal regularizer and a curriculum learning strategy are introduced for model training. The former aims to generate score-discriminative features by exploiting the ordinal nature of regression targets, while the latter gradually ramps up the training complexity to facilitate the assessment task that takes unscripted speech as input. Experiments conducted on the Speechocean762 benchmark dataset validates the feasibility and superiority of our method in relation to several cutting-edge baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。