arXiv:2505.24304eess.AScs.SD2025-05中稿 · Interspeech 2025被引 1

用真人跟读数据模拟母语者听感,更准评估二语发音可懂度。

A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion

  • 通过跟读数据和序列到序列语音转换,模拟母语者听觉感知
  • 能准确识别导致理解困难的发音片段,与真人判断高度一致
  • 适合需要真实听感评估的外语教学系统

评估二语发音可懂度对计算机辅助语言学习(CALL)至关重要。传统基于自动语音识别(ASR)的方法常关注发音接近母语程度,但可能无法捕捉母语者实际感知的可懂度。本文提出一种新型感知驱动的二语可懂度指标,利用母语者跟读数据,在序列到序列(seq2seq)语音转换框架中构建模型。通过引入对齐机制和声学特征重建,该方法模拟母语者听觉感知,识别出二语发音中可能导致理解困难的片段。客观与主观评估均表明,该方法比传统ASR指标更贴近母语者判断,为全球多语环境下的CALL系统提供了新方向。

原文摘要 · Abstract (English)

Evaluating L2 speech intelligibility is crucial for effective computer-assisted language learning (CALL). Conventional ASR-based methods often focus on native-likeness, which may fail to capture the actual intelligibility perceived by human listeners. In contrast, our work introduces a novel, perception based L2 speech intelligibility indicator that leverages a native rater's shadowing data within a sequence-to-sequence (seq2seq) voice conversion framework. By integrating an alignment mechanism and acoustic feature reconstruction, our approach simulates the auditory perception of native listeners, identifying segments in L2 speech that are likely to cause comprehension difficulties. Both objective and subjective evaluations indicate that our method aligns more closely with native judgments than traditional ASR-based metrics, offering a promising new direction for CALL systems in a global, multilingual contexts.

语音识别二语学习感知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。