arXiv:2508.07229cs.CLcs.LG2025-08中稿 · the Journal of the…被引 1

用深度学习分析英语单词重音,揭示模型如何从真实语音中捕捉重音线索。

How Does a Deep Neural Network Look at Lexical Stress in English Words?

  • 通过CNN从声谱图预测双音节词重音位置,避免了人工标注的最小对立项干扰。
  • 最高达到92%准确率,且在最小对立项上仍能正确区分如PROtest与proTEST。
  • 发现模型主要依赖重读音节的元音频谱特征,尤其是一、二共振峰。

尽管神经网络在语音处理中表现优异,但其决策过程常被视为黑箱。本文研究了深度学习在英语词汇重音识别中的可解释性。我们从朗读和自然对话语音中自动构建了一个包含双音节词的数据集,使用多种卷积神经网络(CNN)架构,基于缺乏最小重音对的声谱图预测重音位置,在保留测试集上最高达到92%准确率。利用层间相关传播(LRP)技术分析发现,模型对保留测试集中的最小重音对(如PROtest vs. proTEST)的判断,主要受重读与非重读音节间信息影响,特别是重读元音的谱特性。此外,分类器也关注整词范围的信息。本文还提出一种特征特异性相关性分析,结果显示最优模型强烈依赖重读元音的一、二共振峰,同时有证据表明音高和第三共振峰亦具贡献。这些结果表明,深度学习能够从自然语料中习得重音的分布式线索,拓展了传统语音学研究依赖高度控制刺激的局限。

原文摘要 · Abstract (English)

Despite their success in speech processing, neural networks often operate as black boxes, prompting the question: what informs their decisions, and how can we interpret them? This work examines this issue in the context of lexical stress. A dataset of English disyllabic words was automatically constructed from read and spontaneous speech. Several Convolutional Neural Network (CNN) architectures were trained to predict stress position from a spectrographic representation of disyllabic words lacking minimal stress pairs (e.g., initial stress WAllet, final stress exTEND), achieving up to 92% accuracy on held-out test data. Layerwise Relevance Propagation (LRP), a technique for neural network interpretability analysis, revealed that predictions for held-out minimal pairs (PROtest vs. proTEST ) were most strongly influenced by information in stressed versus unstressed syllables, particularly the spectral properties of stressed vowels. However, the classifiers also attended to information throughout the word. A feature-specific relevance analysis is proposed, and its results suggest that our best-performing classifier is strongly influenced by the stressed vowel's first and second formants, with some evidence that its pitch and third formant also contribute. These results reveal deep learning's ability to acquire distributed cues to stress from naturally occurring data, extending traditional phonetic work based around highly controlled stimuli.

语音识别深度学习可解释性重音预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。