研究语音情绪与语义情感不一致时大模型的识别缺陷
When Vocal Tone and Literal Meaning Diverge: An Acoustic-Semantic Incongruity Study for Large Audio-Language Models

- 构建新数据集CREMA-ASIS,分离声学情绪与语义情感
- 大模型在情绪与语义矛盾时识别率低,主要依赖语义
- 微调后模型在跨域数据上同时提升声学与语义理解
多模态情感线索可能存在不一致(如讽刺或嘲笑式赞美),单靠某一模态易导致误解。尽管大音频语言模型(LALMs)近年在多模态情感识别中广泛应用,其对声学与语义线索的解耦能力,尤其是在矛盾情形下的表现仍缺乏研究。为此,我们提出了CREMA-ASIS数据集,专门用于研究声学情绪与语义情感极性之间的不一致。该数据集将声学情绪标签与语义情感极性进行配对。利用此数据集,我们在多任务框架下评估了LALM的偏见,并通过逐层分析识别各层级的模态主导性。结果表明,LALMs在语义-声学矛盾情况下表现不佳,很少预测出不一致;且整体受语义信息主导。然而,经过监督微调后,模型在我们的CREMA-ASIS测试集上的性能显著提升,同时保持了转录准确性和联合情感识别能力。结果表明,模型在域外数据上具备提升声学与语义理解的潜力。
原文摘要 · Abstract (English)
Affective cues across modalities may be incongruous (e.g., sarcasm or mocking praise), potentially leading to misinterpretation when relying on a single modality. Large Audio-Language Models (LALMs) have recently gained popularity and been applied to multimodal emotion recognition, but their ability to disentangle acoustic and semantic cues, especially in incongruent cases, remains underexplored. To address this gap, we introduce CREMA-ASIS, a dataset specifically created to investigate incongruence between acoustic emotion and semantic sentiment cues. It pairs acoustic emotion labels with semantic sentiment polarities. Using this dataset, we evaluate LALM biases within a multitask framework and conduct a layer-wise analysis to identify modality dominance across layers. Our findings reveal that LALMs struggle with semantic-acoustic incongruent cases, rarely predicting incongruity, and that LALMs are predominantly influenced by semantic information. However, supervised fine-tuning significantly improves LALM performance on our CREMA-ASIS test set while preserving transcription accuracy and joint emotion recognition. Results demonstrate potential for enhancing both acoustic and semantic understanding on out-of-domain data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。