MuFFIN统一建模发音诊断与评估,提升语言学习反馈精度
MuFFIN: Multifaceted Pronunciation Feedback Model with Interactive Hierarchical Neural Modeling
- 构建交互式层级神经网络,联合优化发音诊断与评估
- 引入音素对比排序正则化,增强音素特征区分度
- 设计针对性训练目标缓解发音错误数据不平衡问题
计算机辅助发音训练(CAPT)通过即时反馈帮助第二语言学习者练习发音。现有方法主要分为发音错误检测与诊断(MDD)和自动发音评估(APA)两类,前者定位语音错误并提供诊断,后者量化发音能力的多维度表现。尽管两者具有天然互补性,但研究者常将其视为独立任务,采用不同建模方式。为此,本文提出MuFFIN——一种基于交互式层级神经架构的多维发音反馈模型,联合解决MDD与APA任务。为更好捕捉音素间细微差异,提出音素对比排序正则化机制,优化模型生成更具音素区分性的特征,同时考虑评分的序数特性。针对MDD中的数据不平衡问题,设计一种简单有效的训练目标,通过音素特异性扰动改进分类器输出分布,兼顾误发音特征。在Speechocean762基准数据集上的实验表明,该方法在APA和MDD任务上均达到领先性能。
原文摘要 · Abstract (English)
Computer-assisted pronunciation training (CAPT) manages to facilitate second-language (L2) learners to practice pronunciation skills by offering timely and instructive feedback. To examine pronunciation proficiency from multiple facets, existing methods for CAPT broadly fall into two categories: mispronunciation detection and diagnosis (MDD) as well as automatic pronunciation assessment (APA). The former aims to pinpoint phonetic pronunciation errors and provide diagnostic feedback, while the latter seeks instead to quantify pronunciation proficiency pertaining to various aspects. Despite the natural complementarity between MDD and APA, researchers and practitioners, however, often treat them as independent tasks with disparate modeling paradigms. In light of this, we in this paper first introduce MuFFIN, a Multi-Faceted pronunciation Feedback model with an Interactive hierarchical Neural architecture, to jointly address the tasks of MDD and APA. To better capture the nuanced distinctions between phonemes in the feature space, a novel phoneme-contrastive ordinal regularization mechanism is then put forward to optimize the proposed model to generate more phoneme-discriminative features while factoring in the ordinality of the aspect scores. In addition, to address the intricate data imbalance problem in MDD, we design a simple yet effective training objective, which is specifically tailored to perturb the outputs of a phoneme classifier with the phoneme-specific variations, so as to better render the distribution of predicted phonemes meanwhile considering their mispronunciation characteristics. A series of experiments conducted on the Speechocean762 benchmark dataset demonstrates the efficacy of our method in relation to several cutting-edge baselines, showing state-of-the-art performance on both the APA and MDD tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。