用非语言声音监督语音情感识别,提升多语言低资源场景效果
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition

- 以非语言发声作为监督信号,实现跨语言情感迁移
- 在多语言数据上达到领先性能,显著优于传统方法
- 适合低资源多语言情感分析研究者参考
本文提出一种针对低资源多语言语音情感识别(LRM-SER)的声调辅助监督范式,利用非语言发声中的韵律线索来捕捉情感特征。不同于依赖标注语音语义的传统方法,本工作将LRM-SER重构为非语言到语言的迁移任务:从标注的非语言源域中提取情感原型,适配至多个目标语言的未标注语音。为此,我们提出NOVA ARC——一个基于双曲几何的框架,其在庞加莱球空间建模情感结构,通过超曲向量量化声调代码本离散化副语言模式,并借助超曲情感透镜捕捉情感强度。对于无监督适配,NOVA-ARC基于最优传输进行源情感原型与目标语音之间的原型对齐,生成软标签并辅以一致性正则化稳定训练。实验表明,该方法在非语言到语言迁移及互补的语言到语言迁移设置下均表现最佳,持续超越欧氏基线和强健的自监督学习模型。据我们所知,这是首个突破语音语义中心监督范式的非语言到语言迁移情感识别工作。
原文摘要 · Abstract (English)
In this work, we introduce a paralinguistic supervision paradigm for low-resource multilingual speech emotion recognition (LRM-SER) that leverages non-verbal vocalizations to exploit prosody-centric emotion cues. Unlike conventional SER systems that rely heavily on labeled verbal speech and suffer from poor cross-lingual transfer, our approach reformulates LRM-SER as non-verbal-to-verbal transfer, where supervision from a labeled non-verbal source domain is adapted to unlabeled verbal speech across multiple target languages. To this end, we propose NOVA ARC, a geometry-aware framework that models affective structure in the Poincaré ball, discretizes paralinguistic patterns via a hyperbolic vector-quantized prosody codebook, and captures emotion intensity through a hyperbolic emotion lens. For unsupervised adaptation, NOVA-ARC performs optimal transport based prototype alignment between source emotion prototypes and target utterances, inducing soft supervision for unlabeled speech while being stabilized through consistency regularization. Experiments show that NOVA-ARC delivers the strongest performance under both non-verbal-to-verbal adaptation and the complementary verbal-to-verbal transfer setting, consistently outperforming Euclidean counterparts and strong SSL baselines. To the best of our knowledge, this work is the first to move beyond verbal-speech-centric supervision by introducing a non-verbal-to-verbal transfer paradigm for SER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。