用提示增强与零样本融合,提升语音情感识别性能
Prompt Amplification and Zero-Shot Late Fusion in Audio-Language Models for Speech Emotion Recognition
- 将双编码器音频语言模型与专用基础模型做零样本后期融合
- 通过提示增强重复查询,显著提升零样本情感识别效果
- 在三个数据集上超越现有最优基准,适合多模态情感分析研究者
音频-语言模型(ALMs)在理解语音与非语音音频方面取得进展。然而,针对闭合式语音处理任务如语音情感识别(SER),领域专用基础模型(FMs)仍是最佳选择。使用ALMs进行零样本SER虽常见,但其与专业模型结合以达到最先进(SOTA)性能的潜力尚未探索。本文提出ZS-Fuse,一种后期融合方法,将双编码器ALM的零样本情感估计与专用FM相结合。为应对情感模糊性和对提示选择的敏感性,我们采用简单提示集成,并提出一种新技巧——提示增强:重复音频与文本查询以发掘更强的零样本能力。我们在三个双编码器ALM和两个FM上评估ZS-Fuse,结果表明其在三个语音情感识别数据集上优于现有最先进基线,包括WavLM-Large。
原文摘要 · Abstract (English)
Audio-Language Models (ALMs) are making strides in understanding speech and non-speech audio. However, domain-specialist Foundation Models (FMs) remain the best for closed-ended speech processing tasks such as Speech Emotion Recognition (SER). Using ALMs for Zero-shot SER is a popular choice, but their potential to work with specialists to achieve state-of-the-art (SOTA) performance remains unexplored. We propose ZS-Fuse, a late-fusion method that combines zero-shot emotion estimates from a dual-encoder ALM with specialist FMs. To handle ambiguity in emotions and sensitivity to prompt choice, 1) we use a simple prompt ensemble and 2) suggest a novel technique called prompt amplification, which repeats audio and text queries to discover stronger zero-shot capabilities. We demonstrate the efficacy of our technique by evaluating ZS-Fuse with three dual-encoder ALMs and two FMs, and report improvements over SOTA baselines, such as WavLM-Large, on three speech emotion recognition datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。