arXiv:2506.06820cs.CLcs.SD2025-06被引 3

让语音情绪识别从分类转向有依据的推理,提升准确率与解释性。

Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs

  • 用生成式推理替代传统分类,让模型说出判断情绪的理由。
  • 在IEMOCAP和MELD上准确率提升,生成回答更连贯、有证据支持。
  • 适合需要可解释情绪分析的应用,如智能客服、心理辅助。

音频大语言模型(AudioLLMs)在语音识别、翻译等语义任务中表现优异,但在建模语气、情绪等副语言信息方面仍显不足。现有方法通常将情绪理解视为分类任务,缺乏对预测依据的阐释。本文探索情绪推理策略,利用AudioLLMs的生成能力,通过生成语义一致、基于证据的解释来增强情绪识别。为此,提出统一框架,结合推理增强的数据监督、双编码器结构与任务交替训练,使AudioLLMs在多任务学习中有效融合情感推理能力。在IEMOCAP和MELD数据集上的实验表明,该方法不仅提升情绪预测准确率,还增强生成响应的连贯性与证据支撑度。在两个跨域数据集上的测试验证了模型的泛化能力。

原文摘要 · Abstract (English)

Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues such as emotion. Existing approaches often treat emotion understanding as a classification problem, offering little insight into the underlying rationale behind predictions. In this work, we explore emotion reasoning, a strategy that leverages the generative capabilities of AudioLLMs to enhance emotion recognition by producing semantically aligned, evidence-grounded explanations. To support this in multitask AudioLLMs, we introduce a unified framework combining reasoning-augmented data supervision, dual-encoder architecture, and task-alternating training. This approach enables AudioLLMs to effectively learn different tasks while incorporating emotional reasoning. Experiments on IEMOCAP and MELD show that our approach not only improves emotion prediction accuracy but also enhances the coherence and evidential grounding of the generated responses. Experiments on two out-of-domain datasets demonstrate the generalization capabilities of the resulting model.

语音情绪生成推理多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。