用AI分析德国大选期间Instagram上的政治人物视觉传播,效果优于传统模型。
Seeing Candidates at Scale: Multimodal LLMs for Visual Political Communication on Instagram

- 对比传统视觉模型与多模态大模型,用GPT-4o识别政治人物和人数。
- 在故事中人物识别F1达0.89,人数统计达0.86,表现最优。
- 适合关注社交媒体政治传播分析的研究者参考。
本文开展了一项计算案例研究,评估专用机器学习模型与新兴多模态大语言模型在视觉政治传播(VPC)分析中的能力。聚焦2021年德国联邦选举期间Instagram动态和帖子中的人物集中曝光现象,我们比较了传统计算机视觉模型(FaceNet512、RetinaFace、Google Cloud Vision)与多模态大语言模型(GPT-4o)在识别领先候选人及图像中人物计数方面的表现。GPT-4o在人物识别上达到0.89的宏平均F1分数,在人物计数上达到0.86,显著优于其他模型。研究结果表明,先进AI系统具备规模化与精细化分析政治传播视觉内容的潜力,同时提示未来研究需关注方法论问题。
原文摘要 · Abstract (English)
This paper presents a computational case study that evaluates the capabilities of specialized machine learning models and emerging multimodal large language models for Visual Political Communication (VPC) analysis. Focusing on concentrated visibility in Instagram stories and posts during the 2021 German federal election campaign, we compare the performance of traditional computer vision models (FaceNet512, RetinaFace, Google Cloud Vision) with a multimodal large language model (GPT-4o) in identifying front-runner politicians and counting individuals in images. GPT-4o outperformed the other models, achieving a macro F1-score of 0.89 for face recognition and 0.86 for person counting in stories. These findings demonstrate the potential of advanced AI systems to scale and refine visual content analysis in political communication while highlighting methodological considerations for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。