arXiv:2507.14346eess.AScs.SD2025-07被引 4

通过音素相似性建模,提升发音错误检测准确率。

Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling

  • 构建多任务框架,用音素相似性捕捉真实发音差异。
  • 在模拟数据集上实现更高错误检出率,超越现有方法。
  • 适合语音评估、语言学习系统开发者参考。

发音错误检测是自动发音评估的核心任务,旨在识别音素级别的发音偏差。口音和言语不流畅带来的语音变异使音素识别困难,当前模型难以有效捕捉这些差异。本文提出一种逐字音素识别框架,采用新型音素相似性建模的多任务训练策略,转录说话人实际说出的内容,而非预期内容。我们构建并开源了 extit{VCTK-accent} 模拟数据集,包含音素错误,并提出两种新指标用于评估发音差异。本工作建立了发音错误检测的新基准。

原文摘要 · Abstract (English)

Phonetic error detection, a core subtask of automatic pronunciation assessment, identifies pronunciation deviations at the phoneme level. Speech variability from accents and dysfluencies challenges accurate phoneme recognition, with current models failing to capture these discrepancies effectively. We propose a verbatim phoneme recognition framework using multi-task training with novel phoneme similarity modeling that transcribes what speakers actually say rather than what they're supposed to say. We develop and open-source \textit{VCTK-accent}, a simulated dataset containing phonetic errors, and propose two novel metrics for assessing pronunciation differences. Our work establishes a new benchmark for phonetic error detection.

发音评估音素识别多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。