构建首个多语言语音情绪识别基准,评估大模型在真实情感模糊场景下的表现。
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
- 设计从直接分类到语调推理的多级提示体系
- 引入软标签与提示集成策略,模拟人类标注者分歧
- 发现大模型虽准确率低,但更贴近人对情绪分布的主观感知
语音大语言模型(Speech LLMs)通过生成式接口在语音情绪识别(SER)中展现出巨大潜力。然而,从封闭集分类转向开放文本生成会引入零样本随机性,使评估高度依赖提示设计。此外,传统语音LLM基准忽略了人类情绪的固有模糊性。为此,我们提出VoxEmo,一个涵盖35个情绪语料库、覆盖15种语言的综合性语音情绪识别基准。VoxEmo提供标准化工具包,包含从直接分类到副语言推理的多种提示复杂度。为反映真实感知与应用,我们引入分布感知的软标签协议和提示集成策略,模拟标注者分歧。实验表明,尽管零样本语音LLMs在硬标签准确率上落后于监督基线,但其预测结果更符合人类主观情绪分布。
原文摘要 · Abstract (English)
Speech Large Language Models (LLMs) show great promise for speech emotion recognition (SER) via generative interfaces. However, shifting from closed-set classification to open text generation introduces zero-shot stochasticity, making evaluation highly sensitive to prompts. Additionally, conventional speech LLMs benchmarks overlook the inherent ambiguity of human emotion. Hence, we present VoxEmo, a comprehensive SER benchmark encompassing 35 emotion corpora across 15 languages for Speech LLMs. VoxEmo provides a standardized toolkit featuring varying prompt complexities, from direct classification to paralinguistic reasoning. To reflect real-world perception/application, we introduce a distribution-aware soft-label protocol and a prompt-ensemble strategy that emulates annotator disagreement. Experiments reveal that while zero-shot speech LLMs trail supervised baselines in hard-label accuracy, they uniquely align with human subjective distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。