arXiv:2412.08306eess.AScs.SD2024-12被引 1

比较生成与判别式语音增强模型对音节重音的保护效果,发现生成模型在特定特征下表现更优。

Evaluating the Impact of Discriminative and Generative E2E Speech Enhancement Models on Syllable Stress Preservation

  • 用生成与判别两类语音增强模型处理不同信噪比噪声下的语音数据。
  • 使用启发式特征时,生成模型在0~20 dB信噪比下均保持良好重音识别性能。
  • 真人听觉实验验证了模型结果,说明生成模型更贴近人类感知的重音模式。

自动音节重音检测是计算机辅助语言学习系统中帮助语言学习者的重要组件。现有重音检测模型通常在干净语音上训练,难以适应真实场景中的噪声环境。为此,本文研究语音增强(SE)模型对音节重音模式的保留影响,对比了判别式与生成式建模范式。实验在0至20 dB不同信噪比条件下,对非德语和意大利语母语者的英语语音数据进行测试,并评估不同特征集在噪声中捕捉重音模式的有效性。进一步开展基于人类感知的实验,比较增强语音与原始干净语音的重音感知差异。结果表明,在使用启发式特征时,生成式语音增强模型能有效保持重音检测性能;且感知实验结果与客观检测结果一致,验证了其在重音保真度上的优势。

原文摘要 · Abstract (English)

Automatic syllable stress detection is a crucial component in Computer-Assisted Language Learning (CALL) systems for language learners. Current stress detection models are typically trained on clean speech, which may not be robust in real-world scenarios where background noise is prevalent. To address this, speech enhancement (SE) models, designed to enhance speech by removing noise, might be employed, but their impact on preserving syllable stress patterns is not well studied. This study examines how different SE models, representing discriminative and generative modeling approaches, affect syllable stress detection under noisy conditions. We assess these models by applying them to speech data with varying signal-to-noise ratios (SNRs) from 0 to 20 dB, and evaluating their effectiveness in maintaining stress patterns. Additionally, we explore different feature sets to determine which ones are most effective for capturing stress patterns amidst noise. To further understand the impact of SE models, a human-based perceptual study is conducted to compare the perceived stress patterns in SE-enhanced speech with those in clean speech, providing insights into how well these models preserve syllable stress as perceived by listeners. Experiments are performed on English speech data from non-native speakers of German and Italian. And the results reveal that the stress detection performance is robust with the generative SE models when heuristic features are used. Also, the observations from the perceptual study are consistent with the stress detection outcomes under all SE models.

语音增强重音识别生成模型听觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。