arXiv:2502.19387cs.LGcs.CL2025-02被引 1

通过残差法分离语音音调与语言内容,提升情感分析精度

Residual Speech Embeddings for Tone Classification: Removing Linguistic Content to Enhance Paralinguistic Analysis

  • 用文本嵌入回归语音嵌入,取残差作为纯音调表示
  • 残差嵌入使音调分类准确率显著提升,线性可分性增强
  • 适合做情感分析、说话人识别等非语言特征研究

自监督语音模型(如wav2vec2、HuBERT、WavLM、Whisper)生成的嵌入同时包含语言和副语言信息,难以独立分析音调。本文提出一种方法:将语音嵌入对齐其对应文本嵌入进行回归,使用残差作为语音音调的表示。在多个自监督语音嵌入上评估该方法,结果表明残差嵌入相比原始语音嵌入显著提升音调分类性能。实验显示该方法增强了线性可分性,即使使用逻辑回归等简单模型也能实现更好分类。残差嵌入的可视化证实了语言信息被有效移除,而音调相关特征得以保留。研究结果表明,残差嵌入在情感分析、说话人表征及副语言语音处理中具有应用潜力。

原文摘要 · Abstract (English)

Self-supervised learning models for speech processing, such as wav2vec2, HuBERT, WavLM, and Whisper, generate embeddings that capture both linguistic and paralinguistic information, making it challenging to analyze tone independently of spoken content. In this work, we introduce a method for disentangling paralinguistic features from linguistic content by regressing speech embeddings onto their corresponding text embeddings and using the residuals as a representation of vocal tone. We evaluate this approach across multiple self-supervised speech embeddings, demonstrating that residual embeddings significantly improve tone classification performance compared to raw speech embeddings. Our results show that this method enhances linear separability, enabling improved classification even with simple models such as logistic regression. Visualization of the residual embeddings further confirms the successful removal of linguistic information while preserving tone-related features. These findings highlight the potential of residual embeddings for applications in sentiment analysis, speaker characterization, and paralinguistic speech processing.

语音分析音调识别嵌入解耦副语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。