arXiv:2506.03403eess.AS2025-06中稿 · INTERSPEECH 2025被引 3

将语音情绪识别中的两种不同表示映射到双曲空间融合,性能超越单一或同类融合方法。

HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition

  • 将语音表征从欧氏空间转至双曲空间,实现异质表示的对齐与融合。
  • 在x-vector(RLR)与Soundstream(CBR)融合后达到当前最优(SOTA)性能。
  • 适合关注语音情感识别中多模态表征融合的研究者和工程师。

基于压缩的表示(CBRs)如EnCodec可捕捉音高、音色等精细声学特征,而基于表示学习的表示(RLRs)如WavLM则编码高层语义与韵律信息。以往语音情绪识别(SER)研究分别探索过这两类表征,但尚未系统研究其融合。本文填补这一空白,提出HYFuse框架,通过将两者映射至双曲空间实现融合,并假设二者能提供互补信息。实验表明,融合x-vector(RLR)与Soundstream(CBR)后,性能显著优于单独使用任一表示或同类型表示融合,达到当前最优(SOTA)水平。

原文摘要 · Abstract (English)

Compression-based representations (CBRs) from neural audio codecs such as EnCodec capture intricate acoustic features like pitch and timbre, while representation-learning-based representations (RLRs) from pre-trained models trained for speech representation learning such as WavLM encode high-level semantic and prosodic information. Previous research on Speech Emotion Recognition (SER) has explored both, however, fusion of CBRs and RLRs haven't been explored yet. In this study, we solve this gap and investigate the fusion of RLRs and CBRs and hypothesize they will be more effective by providing complementary information. To this end, we propose, HYFuse, a novel framework that fuses the representations by transforming them to hyperbolic space. With HYFuse, through fusion of x-vector (RLR) and Soundstream (CBR), we achieve the top performance in comparison to individual representations as well as the homogeneous fusion of RLRs and CBRs and report SOTA.

语音识别情绪识别双曲空间表征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。