arXiv:2504.13765eess.AScs.SD2025-04被引 2

用梅尔倒谱系数分析母语对二语发音的影响,让机器可解释且能指导教学。

Modeling L1 Influence on L2 Pronunciation: An MFCC-Based Framework for Explainable Machine Learning and Pedagogical Feedback

  • 基于MFCC特征提取,构建可解释的二语发音差异分析框架。
  • 三个核心MFCC特征使模型准确率显著优于全特征模型。
  • 成果可用于英语教学改进与语音评估工具设计,适合语言教育者。

本研究探讨梅尔频率倒谱系数(MFCC)在捕捉汉语和美式英语母语者在二语英语发音中母语迁移现象方面的有效性。从乔治梅森大学语音口音档案中提取母语为汉语和美式英语的说话人语音样本,转换为WAV格式,并计算每位说话人的13个MFCC特征。通过结合推断统计(t检验、多元方差分析、典型判别分析)与机器学习(随机森林分类)的多方法分析框架,发现MFCC-1(宽带能量)、MFCC-2(第一共振峰区域)和MFCC-5(浊音与摩擦音能量)是最具区分性的特征。使用这三个特征的简化模型在性能上显著优于全特征模型,经麦克内马尔检验及置信区间不重叠验证。研究结果实证支持感知同化模型(PAM-L2)与语音习得模型(SLM),表明母语影响下的二语发音变异既具有感知基础也具备声学可量化性。方法上,为应用语言学与可解释人工智能提供透明、数据高效的二语发音建模流程。研究成果也为英语作为第二语言/外语教学提供了教学启示,明确了可提升发音清晰度的母语特异性特征,适用于智能教学系统、课程设计与语音评估工具开发。

原文摘要 · Abstract (English)

This study investigates the extent to which Mel-Frequency Cepstral Coefficients (MFCCs) capture first language (L1) transfer in extended second language (L2) English speech. Speech samples from Mandarin and American English L1 speakers were extracted from the GMU Speech Accent Archive, converted to WAV format, and processed to obtain thirteen MFCCs per speaker. A multi-method analytic framework combining inferential statistics (t-tests, MANOVA, Canonical Discriminant Analysis) and machine learning (Random Forest classification) identified MFCC-1 (broadband energy), MFCC-2 (first formant region), and MFCC-5 (voicing and fricative energy) as the most discriminative features for distinguishing L1 backgrounds. A reduced-feature model using these MFCCs significantly outperformed the full-feature model, as confirmed by McNemar's test and non-overlapping confidence intervals. The findings empirically support the Perceptual Assimilation Model for L2 (PAM-L2) and the Speech Learning Model (SLM), demonstrating that L1-conditioned variation in L2 speech is both perceptually grounded and acoustically quantifiable. Methodologically, the study contributes to applied linguistics and explainable AI by proposing a transparent, data-efficient pipeline for L2 pronunciation modeling. The results also offer pedagogical implications for ESL/EFL instruction by highlighting L1-specific features that can inform intelligibility-oriented instruction, curriculum design, and speech assessment tools.

语音分析可解释AI语言教学机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。