用多模态融合与大模型结合提升语音情感识别准确率
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
- 设计EmoQ-Former融合多模态信息生成查询向量
- 在IEMOCAP和MELD数据集上达到当前最优性能
- 适合关注语音情感分析与大模型融合的研究者
语音情感识别(SER)受限于单模态系统情感信息不足和多模态系统特征对齐困难。近年来,多模态大语言模型(MLLM)在SER上取得进展,但仍存在复杂情感推理中的幻觉和误分类问题。为此,本文提出基于MLLM的EmoQ框架,通过EmoQ-Former融合多模态信息生成查询嵌入,并采用多目标情感学习(MAL)实现协同优化。框架还引入软提示注入策略,将多模态表示注入大语言模型。该端到端架构在IEMOCAP和MELD数据集上取得当前最优表现,为SER提供了新的多模态融合范式。
原文摘要 · Abstract (English)
The performance of speech emotion recognition (SER) is limited by the insufficient emotion information in unimodal systems and the feature alignment difficulties in multimodal systems. Recently, multimodal large language models (MLLMs) have made progress in SER. However, MLLMs still suffer from hallucination and misclassification problems in complex emotion reasoning. To address these problems, we propose an MLLM-based framework called EmoQ, which generates query embeddings that fuse multimodal information through an EmoQ-Former and uses multi-objective affective learning (MAL) to achieve co-optimization. The framework also provides a soft-prompt injection strategy to inject multimodal representations into the LLM. This end-to-end architecture achieves state-of-the-art performance on the IEMOCAP and MELD datasets, providing a new multimodal fusion paradigm for SER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。