arXiv:2606.24082eess.AScs.SD2026-06

用推理引导的音频模型,仅用5%数据就能更好判断语音情绪差异。

Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions

论文配图:Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions
图 1 · 摘自论文原文
  • 用配对语音输入+推理轨迹训练音频语言模型
  • 在情绪判断上提升预测准确率,仅需传统方法5%数据
  • 适合需要小样本情绪分析的智能系统开发者

大型音频语言模型(LALMs)能处理音频推理,但尚不清楚它们是否能在情绪、环境、语言、语调和人际维度上对两段语音进行比较判断。本文聚焦语音情绪识别(SER),研究模型如何判断哪段语音表现出更高唤醒度、效价或支配性。提出一种基于推理引导的序数级SER框架,将模型置于双语音输入条件下,利用语义音频描述与来自GeMAPS特征的声学证据生成推理轨迹进行训练,实现可解释的比较决策。除直接监督外,还采用直接偏好优化,强化情绪差异的判别能力。实验表明,该框架在偏好预测上表现更优,且仅需传统序数级SER系统5%的训练数据。

原文摘要 · Abstract (English)

Large audio-language models (LALMs) can reason about audio, yet it remains unclear whether they can perform comparative judgments between two speech signals along emotional, environmental, linguistic, prosodic, and interpersonal dimensions. We study this question in the context of speech emotion recognition (SER), where the model determines which utterance exhibits higher arousal, valence, or dominance. We introduce a reasoning-guided ordinal SER framework that conditions an LALM on paired speech inputs. The model is trained using reasoning traces generated from both semantic audio descriptions and acoustic evidence derived from GeMAPS features, enabling interpretable comparative decisions. Beyond direct supervision, we also employ direct preference optimization to encourage stronger separation for emotional differences. Experiments show that the proposed framework improves preference prediction while requiring only 5% of the training data used by conventional ordinal SER systems.

情绪识别音频模型小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。