arXiv:2512.10882cs.CL2025-12被引 2

评测多模态大模型在政治视频情绪分析中的表现,发现实验室数据表现好但真实场景效果差。

Computational emotion analysis with multimodal LLMs: Current evidence on an emerging methodological opportunity

  • 用真实政治辩论和实验室录音对比测试多模态大模型的情绪识别能力。
  • 模型在真实议会辩论中与人工评分相关性仅达中等水平,且普遍低估男性情绪。
  • 揭示当前mLLMs在现实政治语境下的可靠性问题,适合政策分析与模型评估研究者参考。

研究越来越多地利用音视频材料分析政治传播中的情绪。多模态大语言模型(mLLMs)通过上下文学习有望实现此类分析,但目前缺乏系统证据证明其在真实政治场景中是否可靠。本文通过两个互补的人工标注数据集——实验室条件下录制的演讲者表演视频与真实议会辩论视频——评估截至2026年初可获取的开源与闭源权重mLLMs在视频情绪唤醒度测量中的表现。结果揭示显著的实验室与真实场景性能差距:在实验室环境下,模型情绪评分接近人类水平;但在议会辩论中,所有模型的情绪评分与平均人工评分的相关性最多仅为中等。此外,在两个数据集中,除一个外所有模型均表现出系统性性别偏差,对男性发言者的情绪低估更严重,导致整体情绪强度出现净正向偏差。这些发现揭示了当前mLLMs在真实政治视频分析中的重要局限,并建立了一个严谨的评估框架以追踪未来进展。

原文摘要 · Abstract (English)

Research increasingly leverages audio-visual materials to analyze emotions in political communication. Multimodal large language models (mLLMs) promise to enable such analyses through in-context learning. However, we lack systematic evidence on whether current mLLMs can reliably measure emotions in real-world political settings. This paper closes this gap by evaluating open- and closed-weights mLLMs available as of early 2026 in video-based emotional arousal measurement using two complementary human-labeled datasets: speech actor recordings created under laboratory conditions and real-world parliamentary debates. I find a critical lab-vs-field performance gap. In videos created under laboratory conditions, the examined mLLMs arousal scores approach human-level reliability. However, in parliamentary debate recordings, all examined models' arousal scores correlate at best moderately with average human ratings. Moreover, in each dataset, all but one of the examined mLLMs exhibit systematic gender-differential bias, consistently underestimating arousal more for male than for female speakers, resulting in a net-positive intensity bias. These findings reveal important limitations of current mLLMs for real-world political video analysis and establish a rigorous evaluation framework for tracking future developments.

情绪分析多模态模型政治传播偏差检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。