通过多任务预训练提升二语发音评估的可解释性与全面性
Multi-task Pretraining for Enhancing Interpretable L2 Pronunciation Assessment
- 采用掩码重建策略捕捉语音长期时序特征
- 在Speechocean762上实现发音评分提升与口语评估相关性增强
- 融合人工设计特征,支持流畅度与重音等可解释评估
自动发音评估(APA)通过细粒度反馈分析第二语言学习者的语音表现。现有方法多依赖音段级特征,导致跨语言层级(如音素、单词、语句)评估仅基于音素级特征,忽略超音段线索。为此,本文提出多任务预训练(MTP)策略:对发音模型的音素级编码器,随机掩码音段级特征并基于上下文重建。同时,当前系统缺乏与自动口语评估(ASA)的整合,限制了综合能力评估。基于经验研究与先验知识,框架引入人工设计特征(HCFs),如流畅度(语速、停顿时长)和重音(语调强度),通过回归器生成可解释的综合评分。在Speechocean762数据集上的实验表明,该方法提升了发音评分精度与ASA相关性,支持精准教学干预与全面能力评估。
原文摘要 · Abstract (English)
Automatic pronunciation assessment (APA) analyzes second-language (L2) learners' speech by providing fine-grained pronunciation feedback at various linguistic levels. Most existing efforts on APA typically adopt segmental-level features as inputs and predict pronunciation scores at different granularities via hierarchical (or parallel) pronunciation modeling. This, however, inevitably causes assessments across linguistic levels (e.g., phone, word, and utterance) to rely solely on phoneme-level pronunciation features, nearly sidelining supra-segmental pronunciation cues. To address this limitation, we introduce multi-task pretraining (MTP) for APA, a simple yet effective strategy that attempts to capture long-term temporal pronunciation cues while strengthening the intrinsic structures within an utterance via the objective of reconstructing input features. Specifically, for a phoneme-level encoder of an APA model, the proposed MTP strategy randomly masks segmental-level pronunciation features and reconstructs the masked ones based on their surrounding pronunciation context. Furthermore, current APA systems lack integration with automated speaking assessment (ASA), limiting holistic proficiency evaluation. Drawing on empirical studies and prior knowledge in ASA, our framework bridges this gap by incorporating handcrafted features (HCFs), such as fluency (speech rate, silence duration) and stress (pitch accent strength), derived from human-designed formulas via regressors to generate interpretable proficiency scores. Experiments on speechocean762 show improved pronunciation scoring and ASA proficiency correlation, enabling targeted training and comprehensive proficiency assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。