arXiv:2607.13721cs.CLeess.AS2026-07

用自监督语音表示和动态时间对齐,无需文本即可评估二语发音、节奏和语调。

Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring

论文配图:Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring
图 1 · 摘自论文原文
  • 基于WavLM的语音表示与动态时间对齐,实现无文本评分框架。
  • 在整体发音评分上超越人类标注者一致率,节奏评估接近人类水平。
  • 适合缺乏标注数据的低资源语言教学场景,尤其关注语音韵律特征。

第二语言语音评估传统上聚焦于音素层面,对节奏、语调等超音段特征的评分研究不足。且多数方法依赖带标签的二语语音训练数据,在低资源环境下难以应用。本文探究利用自监督WavLM表示与动态时间对齐(DTW)是否可构建无需文本的英语和日语二语发音、节奏及语调评估框架。结果表明,仅通过比较学习者语音与母语模板的DTW距离,即可在整体及句子级发音评分上超过人类一致性;针对节奏,提出通过分析DTW对齐路径的扭曲程度进行度量,最优方法已接近人类水平;对于语调,则结合韵律残差的DTW距离与基频、强度特征,但部分任务表现仍较有限。结果表明,自监督语音表示为多维度发音评估提供了有前景的无文本基础。

原文摘要 · Abstract (English)

L2 speech assessment has traditionally focused on phonetic assessment, leaving the scoring of suprasegmental features such as rhythm and intonation underexplored. Moreover, assessment methods often require training with labeled L2 speech data, making them difficult to apply in low-resource settings. We investigate whether DTW over self-supervised WavLM representations can provide a text-free framework for assessing phonetic accuracy, rhythm, and intonation in English and Japanese L2 speech. Results show that a basic DTW-based approach that compares learner speech to native templates exceeds human agreement on holistic and sentence-level phonetic scoring. For rhythm, we introduce methods that measure the degree of warping in the DTW alignment path; our best method approaches human-level performance. For intonation, we combine DTW distance over prosodic residuals with pitch and intensity features, but performance remains more modest on some tasks. Our results point to self-supervised representations as a promising, text-free basis for multi-aspect pronunciation assessment.

语音评估自监督学习二语习得韵律分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。