arXiv:2605.22732cs.AIcs.CL2026-05

用大模型分析政治演讲中的情感,比纯声学方法更准。

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models

  • 结合语音与文本的LLM多模态分析,捕捉政治话语情感。
  • 大模型得出的情感评分与专家评估相关性达0.664(显著)。
  • 适合研究政治话语情感、需语义理解的场景。

我们探究声学情感识别模型能否作为政治演讲中'诉诸情感(Pathos)'维度的有效代理指标,该维度由TRUST多智能体大语言模型(LLM)流程定义。以德国议会菲利克斯·巴纳扎克的一次51段、共245秒的发言为案例,比较三种分析方式:(1) emotion2vec_plus_large,基于声学的语音情感识别(SER)模型,通过后处理的Russell环形模型投影生成连续唤醒度与效价值;(2) Gemini 2.5 Flash,一个能结合音频与转录文本进行开放式上下文感知分析的LLM;(3) TRUST-Pathos得分,来自三位评委型LLM组成的集成系统。斯皮尔曼等级相关分析显示,Gemini的效价与TRUST-Pathos高度相关(rho = +0.664, p < 0.001),而emotion2vec的效价无显著关联(rho = +0.097, p = 0.499)。进一步通过对柏林情绪语音数据库(EMO-DB)采用开放标注范式进行系统质量评估,发现传统SER基准数据集存在表演化语音、文化偏见及类别不兼容问题。结果表明,基于大模型的多模态分析在捕捉语义层面的政治情感方面显著优于仅依赖声学模型的方法,而声学特征仍对低层次唤醒度估计具有价值。未来工作将扩展至融合面部表情与凝视的视频分析。

原文摘要 · Abstract (English)

We investigate whether acoustic emotion recognition models can serve as proxies for the Pathos dimension in political speech analysis, as operationalised by the TRUST multi-agent large language model (LLM) pipeline. Using a Bundestag plenary speech by Felix Banaszak (51 segments, 245 s) as a case study, we compare three analysis modalities: (1) emotion2vec_plus_large, an acoustic speech emotion recognition (SER) model whose continuous Arousal and Valence values are derived via post-hoc Russell Circumplex projection; (2) Gemini 2.5 Flash, an LLM analysing the full speech audio together with its transcript in an open-ended, context-aware fashion; and (3) TRUST-Pathos scores from a three-advocate LLM supervisor ensemble. Spearman rank correlations reveal that Gemini Valence correlates strongly with TRUST-Pathos (rho = +0.664, p < 0.001), whereas emotion2vec Valence does not (rho = +0.097, p = 0.499). We further demonstrate, via a systematic quality evaluation of the Berlin Database of Emotional Speech (EMO-DB) using Gemini in an open-ended annotation paradigm, that standard SER benchmark corpora suffer from acted speech, cultural bias, and category incompatibility. Our results suggest that LLM-based multimodal analysis captures semantically defined political emotion substantially better than acoustic models alone, while acoustic features remain informative for low-level Arousal estimation. Future work will extend this approach to video-based analysis incorporating facial expression and gaze.

情感分析多模态大模型政治演讲

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。