arXiv:2609.04236eess.AScs.SD2026-09

解决语音情绪识别中语义与语调冲突问题,提出新基准和鲁棒框架。

Robust Speech Emotion Recognition under Tone-Word Conflict: A Benchmark and Framework

论文配图:Robust Speech Emotion Recognition under Tone-Word Conflict: A Benchmark and Framework
图 1 · 摘自论文原文
  • 分离声学与语义路径,用注意力机制自适应融合特征
  • 在语调语义冲突场景下性能显著优于现有模型
  • 适合真实复杂对话场景的语音情绪识别研究者

语音情绪识别(SER)是人机交互的关键技术,但现有系统通常假设语调与词汇语义一致,忽略了现实中语调与语义冲突(即情绪表达与字面意思矛盾)的情况。为此,我们提出TWIN-SER基准,用于系统评估声学-语义不一致下的表现,并发现当前先进模型在此类场景下性能严重下降。为应对该问题,我们提出DAS(解耦声学-语义融合)框架,通过显式分离声学与语义路径,选取高能量判别性嵌入,并利用轻量级查询注意力机制自适应融合。DAS包含三个核心模块:异构特征提取模块分别捕获原始输入中的声学与语义表征;高能量嵌入选择模块保留最具判别性的特征;Q-Former组合模块通过交叉注意力连接双路径,实现冲突条件下的鲁棒情绪预测。大量实验表明,DAS在语调语义冲突场景下持续优于现有方法,且在标准域内与零样本设置中表现优异。代码与数据集已开源。

原文摘要 · Abstract (English)

Speech emotion recognition (SER) is a crucial component of human-computer interaction, attracting extensive attention from both industry and academia. However, existing SER systems typically assume alignment between vocal tone and lexical semantics, overlooking the real-world scenarios that involve tone-word conflict-where the emotion conveyed by speech contradicts the literal meaning of the words. To bridge this gap, we introduce TWIN-SER (Tone-Word Incongruent SER), a benchmark for systematic evaluation under acoustic-semantic incongruence, and show that state-of-the-art models degrade severely under such incongruence. To address this, we propose DAS (Disentangled Acoustic-Semantic fusion), a framework that mitigates tone-word conflict by explicitly disentangling acoustic and semantic pathways, selecting informative high-energy embeddings, and adaptively fusing them via a lightweight query-based attention mechanism. Specifically, DAS comprises three crucial modules: i) a heterogeneous feature extraction module that separately captures complementary acoustic and semantic representations from raw input; ii) a high-energy embedding selection module that identifies and retains the most discriminative embeddings; and iii) a Q-Former combination module that bridges the two pathways through cross-attention, enabling robust emotion prediction under incongruent conditions. Extensive experiments demonstrate that DAS consistently outperforms existing methods in tone-word conflict scenarios, as well as in standard in-domain and zero-shot settings. Our code and datasets are available at https://github.com/24DavidHuang/FAS

语音识别情绪识别鲁棒性跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。