无需目标语言标签即可实现跨语言情感识别,仅需少量源语言数据。
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition

- 构建动态情绪-语义结构,通过瞬时共振场引导无标签数据自组织。
- 在多语言实验中仅用5样本标注即达到高精度,超越现有方法。
- 适合低资源语言情感分析,尤其适用于无对齐标注的场景。
跨语言语音情感识别(CLSER)旨在识别未见语言中的情绪状态。现有方法严重依赖完整标签的语义同步和静态特征稳定性,导致低资源语言难以达到高资源语言的性能。为此,我们提出一种基于语义-情感共鸣嵌入(SERE)的半监督框架,这是一种无需目标语言标签或翻译对齐的跨语言动态特征范式。SERE利用少量标注样本构建情绪-语义结构,通过瞬时共振场(IRF)学习人类情感体验,使无标签样本能够自我组织进该结构,实现半监督语义引导与结构发现。此外,设计了三重共鸣交互链(TRIC)损失,增强情感关键期中已标注与未标注样本间的交互与嵌入能力。多语言实验证明该方法有效,仅需源语言5样本标注即可达成优异性能。
原文摘要 · Abstract (English)
Cross-lingual Speech Emotion Recognition (CLSER) aims to identify emotional states in unseen languages. However, existing methods heavily rely on the semantic synchrony of complete labels and static feature stability, hindering low-resource languages from reaching high-resource performance. To address this, we propose a semi-supervised framework based on Semantic-Emotional Resonance Embedding (SERE), a cross-lingual dynamic feature paradigm that requires neither target language labels nor translation alignment. Specifically, SERE constructs an emotion-semantic structure using a small number of labeled samples. It learns human emotional experiences through an Instantaneous Resonance Field (IRF), enabling unlabeled samples to self-organize into this structure. This achieves semi-supervised semantic guidance and structural discovery. Additionally, we design a Triple-Resonance Interaction Chain (TRIC) loss to enable the model to reinforce the interaction and embedding capabilities between labeled and unlabeled samples during emotional highlights. Extensive experiments across multiple languages demonstrate the effectiveness of our method, requiring only 5-shot labeling in the source language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。