融合声学生理机制的语音情绪识别方法,提升模型可解释性与效率。
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
- 基于声带与声道生理结构,构建幅度与相位联合表征
- 在14个数据集、10种语言上实现高精度情绪识别
- 可实时部署于人形机器人,适合需安全交互的场景
语音情绪识别(SER)对人形机器人社交互动和心理诊断等任务至关重要,可解释且高效的模型对安全性与性能尤为关键。现有深度模型在大规模数据上训练,但大多不可解释,未能充分建模情绪相关的声学信号,也未捕捉情绪发声行为的核心生理机制。生理研究显示,声调的幅度与相位动态通过声道滤波器与声门源相关联。然而,多数深度模型仅关注幅度,忽略幅度与相位间的生理耦合。为此,本文提出PhysioSER,一种融合生理知识的语音时频表征学习方法,采用紧凑可插拔设计。该方法基于语音解剖与生理(VAP)知识,构建幅度与相位双视图,以增强自监督学习(SSL)模型。框架包含两条并行路径:一是基于VAP的声学特征表示分支,将信号分解后嵌入四元数空间,使用哈密顿结构四元数卷积建模其动态交互;二是基于冻结的SSL主干的潜在表示分支。随后,通过对比投影与对齐框架对两种路径的句级特征进行对齐,并由浅层注意力融合头完成情绪分类。经14个数据集、10种语言、6种主干网络的广泛评估,证明PhysioSER在可解释性与效率方面表现优异,且已在人形机器人平台上实现实时部署,验证其实际有效性。
原文摘要 · Abstract (English)
Speech emotion recognition (SER) is essential for humanoid robot tasks such as social robotic interactions and robotic psychological diagnosis, where interpretable and efficient models are critical for safety and performance. Existing deep models trained on large datasets remain largely uninterpretable, often insufficiently modeling underlying emotional acoustic signals and failing to capture and analyze the core physiology of emotional vocal behaviors. Physiological research on human voices shows that the dynamics of vocal amplitude and phase correlate with emotions through the vocal tract filter and the glottal source. However, most existing deep models solely involve amplitude but fail to couple the physiological features of and between amplitude and phase. Here, we propose PhysioSER, a physiology-informed vocal spectrotemporal representation learning method, to address these issues with a compact, plug-and-play design. PhysioSER constructs amplitude and phase views informed by voice anatomy and physiology (VAP) to complement SSL models for SER. This VAP-informed framework incorporates two parallel workflows: a vocal feature representation branch to decompose vocal signals based on VAP, embed them into a quaternion field, and use Hamilton-structured quaternion convolutions for modeling their dynamic interactions; and a latent representation branch based on a frozen SSL backbone. Then, utterance-level features from both workflows are aligned by a Contrastive Projection and Alignment framework, followed by a shallow attention fusion head for SER classification. PhysioSER is shown to be interpretable and efficient for SER through extensive evaluations across 14 datasets, 10 languages, and 6 backbones, and its practical efficacy is validated by real-time deployment on a humanoid robotic platform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。