arXiv:2608.28932cs.CLcs.SD2026-08被引 1

测试AI模型从音频中识别情绪的能力,发现表现仍很脆弱。

VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models

  • 构建公开音频情绪评测集,仅用原始音频不依赖文字
  • 六种基线平均准确率35.5%,最强模型仅46.5%
  • 中性情绪识别最准,惊恐与惊讶识别率不足16%

语音产品日益需要捕捉语音中的情感线索,而这些线索在文本转录中缺失。我们提出VocalAffectBench,一个公开的、仅用于测试的基准,用于评估AI音频模型能否从原始音频中识别表达出的情绪。该基准包含273段由51个说话人录制的英语WAV音频片段,总计1.95小时,涵盖七类情绪:愤怒、厌恶、恐惧、快乐、中性、悲伤和惊讶,每类39段。所有基线均仅基于音频进行评估,不使用文本或上下文元数据。在六个已发布的基线中,平均准确率为35.5%。最强基线gemini_3_5_flash在七分类任务中达到46.5%,高于14.3%的随机基线,但远未实现稳健的情绪识别。二级效价分桶分析将标签归为正向、中性、负向三类(排除效价模糊的惊讶),总体准确率为50.9%。各类别表现极不均衡:按召回率计算,中性情绪识别最可靠,平均达75.6%,而惊讶和恐惧的召回率分别仅为10.7%和15.4%。结果表明,当前基线虽能提取部分情感信号,但离散表达情绪识别仍不稳定,尤其对非中性情绪——而这正是语音助手工作流中最关键的部分。

原文摘要 · Abstract (English)

Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-only benchmark for evaluating whether AI audio models can identify expressed vocal emotion from raw audio. The benchmark contains 273 human-recorded English WAV clips from 51 speaker accounts totaling 1.95 hours across seven labels: angry, disgusted, fearful, happy, neutral, sad, and surprised, with 39 clips per class. All baselines are evaluated from audio alone, without transcripts or contextual metadata. Across six released baselines, average accuracy is 35.5%. The strongest baseline, gemini_3_5_flash, reaches 46.5% on the seven-way task, above the 14.3% random baseline but far from robust emotion recognition. A secondary valence-bucket analysis maps labels into positive, neutral, and negative classes, excluding surprised because its valence is ambiguous. Aggregate accuracy under this coarser view is 50.9%. Performance is highly uneven across classes. By recall, neutral is identified most reliably at 75.6% averaged across baselines, while surprised and fearful reach only 10.7% and 15.4%, respectively. These results show that the evaluated baselines can extract some affective signal from speech, but discrete expressed-emotion recognition remains fragile, especially for non-neutral emotions that are often most important in voice agent workflows.

语音情绪识别AI评测音频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。