arXiv:2507.04048cs.SDeess.AS2025-07中稿 · Interspeech2025被引 12

通过软提示微调提升语音情绪识别跨域泛化能力

CLEP-DG: Contrastive Learning for Speech Emotion Domain Generalization via Soft Prompt Tuning

  • 用情感语音数据微调CLAP,增强情绪特征编码
  • 引入文本驱动的声学上下文提示,无需额外标注音频
  • 利用跨模态迁移,在推理时提升对不同环境的适应性

语音情绪识别(SER)是情感计算和人机交互的基础,但现有模型在不同声学条件下的泛化能力有限。尽管对比语言-音频预训练(CLAP)具备强跨模态对齐能力,却缺乏专门捕捉情绪线索的机制,限制了其在SER中的表现。为此,我们提出CLEP-DG框架,增强CLAP在情绪识别中的鲁棒性。首先,基于大规模情感语音数据微调CLAP,得到CLEP,以更好地编码情绪相关特征;其次,提出声学上下文提示微调(ACPT),一种文本驱动的数据增强策略,通过可学习的提示向量建模多样声学环境,无需额外标注音频;最后,利用跨模态可迁移性,使用文本生成的嵌入训练分类器,并在推理时应用于音频编码器,缓解文本监督与音频情绪识别间的域偏移。在五个基准数据集上的实验表明,CLEP-DG优于以往基于CLAP的方法,在有监督和域泛化设置下均达到最先进性能。

原文摘要 · Abstract (English)

Speech Emotion Recognition (SER) is fundamental to affective computing and human-computer interaction, yet existing models struggle to generalize across diverse acoustic conditions. While Contrastive Language-Audio Pretraining (CLAP) provides strong multimodal alignment, it lacks dedicated mechanisms for capturing emotional cues, making it suboptimal for SER. To address this, we propose CLEP-DG, a framework that enhances CLAP's robustness in emotion recognition. First, we fine-tune CLAP to obtain CLEP, adapting it on large-scale emotional speech datasets to better encode emotion-relevant features. Then, we introduce Acoustic Context Prompt Tuning (ACPT), a text-driven augmentation strategy that optimizes learnable prompt vectors to model diverse acoustic environments without additional labeled audio. Finally, leveraging cross-modal transferability, we train a classifier on text-derived embeddings and apply it to the audio encoder during inference, mitigating domain shifts between textual supervision and audio-based emotion recognition. Experiments across five benchmark datasets show that CLEP-DG outperforms prior CLAP-based approaches, achieving state-of-the-art performance in both supervised and domain generalization settings.

语音情绪识别域泛化提示微调跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。