用声学符号提示提升语音情感识别,让模型更懂声音中的情绪线索。
Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition

- 从标准声学特征中提取6个可解释的声学概念符号,加入文本提示
- 对齐符号使识别准确率提升,错乱符号则导致误判向中性偏移
- 符号干预可检测模型是否真正依赖音频信号,适合研究情绪计算可靠性
指令遵循的音频语言模型(ALMs)可引入显式声学线索,但当原始音频已存在时,这些线索是否被合理使用仍不明确。本文在语音情感识别(SER)任务中,从标准化的eGeMAPS副语言特征集中提取6个可解释的声学概念标记:能量、音高、动态、明亮度、共振峰和嗓音质量。这些标记被添加至文本提示,而音频输入保持不变。在FAU-Aibo与IEMOCAP两个常用基准上,对齐的标记显著提升了未加权平均召回率(UAR),而打乱、冲突或损坏的标记则导致性能下降,并使混淆更多指向中性情绪。重要的是,在强符号扰动下模型预测并未崩溃,表明模型对符号通道敏感,但仍部分锚定在音频信号上。我们主张,仅通过符号干预即可有效探测基于ALM的情感计算中线索使用的接地性、鲁棒性与可解释性。
原文摘要 · Abstract (English)
Instruction-following audio language models (ALMs) can be augmented with explicit acoustic cues, yet it remains unclear whether such cues are used in a grounded way when the raw audio is already available. We study this question in speech emotion recognition (SER) by deriving six interpretable acoustic concept tokens from the standardised eGeMAPS paralinguistic feature set. These tokens summarise energy, pitch, dynamics, brightness, formants, and voice quality, and are appended to the textual prompt while the audio input is kept unchanged. Across the widely used FAU-Aibo and IEMOCAP benchmarks, aligned tokens improve unweighted average recall (UAR), whereas shuffled, conflicting, or corrupted tokens reduce performance relative to aligned tokens and shift confusions toward neutral. Importantly, predictions do not collapse under strong token perturbations, suggesting that the models are sensitive to the symbolic cue channel but remain partly anchored to the audio signal. We argue that token-only interventions provide a practical way to probe audio-grounded cue use, robustness, and interpretability in ALM-based affective computing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。