用自监督学习发现语音韵律的深层结构,提升情感识别等长时任务表现。
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
- 设计掩码韵律模型,捕捉语音中非词汇的韵律结构。
- 在情感识别等长时任务上,性能显著优于原始声学特征。
- 揭示了自监督训练时间尺度对复杂结构建模的关键作用。
人们在理解文本时会利用词汇结构的可预测性。尽管语音中也存在可预测的结构,但韵律(如语调、节奏、音量)独立于词汇内容所贡献的结构程度尚不明确。本研究采用自监督学习(SSL)方法,探究韵律声学特征中的时间粒度结构。所提出的掩码韵律模型生成的表示能有效预测依赖局部信息的感知标签(如词边界),但在涉及长期结构的标签(如情感识别)上表现最优。跨多种感知标签的探针实验显示,该模型相较未变换的基线声学特征(基线特征包括音高、能量和语音活动)有显著提升。结果表明,自监督学习训练目标的时间尺度至关重要,且复杂结构比传统受限结构更具价值。
原文摘要 · Abstract (English)
People exploit the predictability of lexical structures during text comprehension. Though predictable structure is also present in speech, the degree to which prosody, e.g. intonation, tempo, and loudness, contributes to such structure independently of the lexical content is unclear. This study leverages self-supervised learning (SSL) to examine the temporal granularity of structures in the acoustic correlates of prosody. Representations from our proposed Masked Prosody Model can predict perceptual labels dependent on local information, such as word boundaries, but provide the most value for labels involving longer-term structures, like emotion recognition. Probing experiments across various perceptual labels show strong relative gains over untransformed pitch, energy, and voice activity features. Our results reveal the importance of SSL training objective timescale and highlight the value of complex SSL-encoded structures compared to more constrained classical structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。