arXiv:2508.11566eess.AScs.CL2025-08中稿 · IEEE ASRU 2025

研究语音模型如何区分重读与非重读词,发现其编码方式具有系统性差异。

Emphasis Sensitivity in Speech Representations

  • 用对比表示差值定义重读,捕捉重读的相对关系。
  • 重读特征与语速变化强相关,但对词义预测效果差。
  • 下游任务微调后重读表征更紧凑,说明其为低维结构化变换。

本文研究现代语音模型是否对语调重读敏感,即是否以系统化方式编码重读词与非重读词。以往工作依赖孤立声学特征(如音高、时长)或标签预测,忽略了重读的相对结构。本文提出基于残差的框架,将重读定义为成对中性与重读词表示的差异。在自监督语音模型上的分析表明,这些残差与时长变化强相关,但在词义预测上表现不佳,说明重读被结构化地以关系形式编码。在语音识别微调模型中,残差占据的子空间比预训练模型紧凑最多50%,进一步表明重读是以一致、低维变换形式编码,且随任务学习变得更加结构化。

原文摘要 · Abstract (English)

This work investigates whether modern speech models are sensitive to prosodic emphasis - whether they encode emphasized and neutral words in systematically different ways. Prior work typically relies on isolated acoustic correlates (e.g., pitch, duration) or label prediction, both of which miss the relational structure of emphasis. This paper proposes a residual-based framework, defining emphasis as the difference between paired neutral and emphasized word representations. Analysis on self-supervised speech models shows that these residuals correlate strongly with duration changes and perform poorly at word identity prediction, indicating a structured, relational encoding of prosodic emphasis. In ASR fine-tuned models, residuals occupy a subspace up to 50% more compact than in pre-trained models, further suggesting that emphasis is encoded as a consistent, low-dimensional transformation that becomes more structured with task-specific learning.

语音表示重读编码自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。