用自监督模型发现语音与音乐情感的声学共性,提升跨领域情感识别效果。
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
- 分析语音与音乐自监督模型各层特征,揭示声学相似性机制。
- 通过两阶段微调,实现语音与音乐情感识别性能双双提升。
- 发现模型存在情绪偏差,但参数高效微调可有效缓解问题。
语音与音乐的情感识别因声学特征重叠而具有相似性,促使研究者关注跨领域知识迁移。然而,语音与音乐中由自监督学习(SSL)模型编码的共享声学线索仍鲜有探索,尤其因语音与音乐的SSL模型在跨域研究中应用较少。本文重新审视情感语音与音乐间的声学相似性,首先分析了语音情感识别(SER)与音乐情感识别(MER)所用的SSL模型的逐层行为;其次,在两阶段微调框架下比较多种方法,探索利用音乐提升SER、用语音提升MER的有效路径;最后,采用弗雷切特音频距离(Frechet audio distance)对各类情绪的声学相似性进行量化分析,揭示了语音与音乐的SSL模型均存在情绪偏差。研究发现,尽管两类模型能捕捉共享声学特征,其表现受训练策略和领域特性影响,不同情绪下差异显著。此外,参数高效微调能有效提升SER与MER性能。本工作为情感语音与音乐的声学相似性提供了新视角,表明跨领域泛化具备潜力以优化情感识别系统。
原文摘要 · Abstract (English)
Emotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and music, particularly those encoded by Self-Supervised Learning (SSL) models, remain largely unexplored, given the fact that SSL models for speech and music have rarely been applied in cross-domain research. In this work, we revisit the acoustic similarity between emotion speech and music, starting with an analysis of the layerwise behavior of SSL models for Speech Emotion Recognition (SER) and Music Emotion Recognition (MER). Furthermore, we perform cross-domain adaptation by comparing several approaches in a two-stage fine-tuning process, examining effective ways to utilize music for SER and speech for MER. Lastly, we explore the acoustic similarities between emotional speech and music using Frechet audio distance for individual emotions, uncovering the issue of emotion bias in both speech and music SSL models. Our findings reveal that while speech and music SSL models do capture shared acoustic features, their behaviors can vary depending on different emotions due to their training strategies and domain-specificities. Additionally, parameter-efficient fine-tuning can enhance SER and MER performance by leveraging knowledge from each other. This study provides new insights into the acoustic similarity between emotional speech and music, and highlights the potential for cross-domain generalization to improve SER and MER systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。