arXiv:2606.14004eess.AS2026-06

提出无监督方法提取纯净的语音韵律嵌入,提升语音分析鲁棒性。

Unsupervised Approaches for Global Prosodic Embedding Extraction

论文配图:Unsupervised Approaches for Global Prosodic Embedding Extraction
图 1 · 摘自论文原文
  • 基于音高和能量的自编码器模型,分离出全局韵律特征。
  • 在复杂场景下性能优于或媲美现有方法,验证了有效性。
  • 适合韵律分析、语音合成等需纯净韵律信息的任务。

韵律是口语交流的核心,传递说话人情绪状态及语义消歧线索。许多自监督语音模型生成的嵌入同时包含韵律、语言和说话人信息,这种混杂在训练与部署条件不一致时会降低模型鲁棒性。若仅关注韵律,则需更纯净的表示。此类表示可用于分析韵律在任务中的作用,或作为语音合成系统的输入。本文提出多种基于音高和能量自编码器的全局韵律嵌入提取方法,并构建基准测试评估其性能。结果表明,在挑战性条件下,所提嵌入相较多种替代方案表现更具竞争力或更优。

原文摘要 · Abstract (English)

Prosody is central to oral communication, conveying information like the emotional state of the speaker and cues needed for meaning disambiguation. Many self-supervised models of speech produce embeddings that encode prosodic as well as linguistic, and speaker information. This entanglement of information is problematic in scenarios where prosody is the main distinguishing factor while other factors may vary between training and deployment; in such cases, a purely prosodic representation would be more robust. Such representation could also be used for analyzing the role of prosody in a given task or as input to speech synthesis systems. In this work, we propose a variety of approaches for producing global prosodic embeddings based on auto-encoder models of pitch and energy. We develop a benchmark for assessing the performance of these representations, showing that our embeddings provide competitive or superior performance under challenging conditions, compared to various alternatives.

韵律建模自编码器语音表示无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。