arXiv:2604.21786cs.CV2026-04

用视觉模型分析社交媒体气候议题,发现大模型能高效捕捉公众讨论趋势。

From Codebooks to VLMs: Evaluating Automated Visual Discourse Analysis for Climate Change on Social Media

论文配图:From Codebooks to VLMs: Evaluating Automated Visual Discourse Analysis for Climate Change on Social Media
图 1 · 摘自论文原文
  • 用提示工程和多模型对比,评估视觉语言模型在气候话语分析中的表现。
  • 大模型Gemini在1038张图像和120万图像上均领先,分布级分析更可靠。
  • 适合做大规模社会议题舆情研究的学者或政策分析人员参考。

社交媒体已成为气候传播的主要阵地,生成了数百万张图片和帖子,若系统分析,可揭示哪些传播策略能引发公众关注,哪些失效。本文旨在推动此类研究,评估计算机视觉方法在社交媒体话语分析中的应用。研究涵盖基于任务的分类体系设计、模型选择、提示工程与验证。我们在两个来自X(原推特)的数据集上基准测试了六种可提示视觉语言模型和十五种零样本CLIP类模型:一个由1,038张专家标注的图像集,另一个包含超过120万张图像、经50,000个标签人工验证的大规模语料库,覆盖五个标注维度:动物内容、气候变化后果、气候行动、图像场景与图像类型。结果显示,Gemini-3.1-flash-lite在所有超类别及两个数据集上均表现最佳,而中等规模开源模型与其差距较小。除个体精度外,我们主张采用分布级评估:即使单图准确率中等,VLM预测仍能可靠还原群体趋势,适合作为大规模话语分析的起点。研究还发现,链式思维推理反而降低性能,而针对具体标注维度设计提示可提升效果。代码与推文ID及标签已公开于https://github.com/KathPra/Codebooks2VLMs.git。

原文摘要 · Abstract (English)

Social media platforms have become primary arenas for climate communication, generating millions of images and posts that - if systematically analysed - can reveal which communication strategies mobilise public concern and which fall flat. We aim to facilitate such research by analysing how computer vision methods can be used for social media discourse analysis. This analysis includes application-based taxonomy design, model selection, prompt engineering, and validation. We benchmark six promptable vision-language models and 15 zero-shot CLIP-like models on two datasets from X (formerly Twitter) - a 1,038-image expert-annotated set and a larger corpus of over 1.2 million images, with 50,000 labels manually validated - spanning five annotation dimensions: animal content, climate change consequences, climate action, image setting, and image type. Among the models benchmarked, Gemini-3.1-flash-lite outperforms all others across all super-categories and both datasets, while the gap to open-weight models of moderate size remains relatively small. Beyond instance-level metrics, we advocate for distributional evaluation: VLM predictions can reliably recover population level trends even when per-image accuracy is moderate, making them a viable starting point for discourse analysis at scale. We find that chain-of-thought reasoning reduces rather than improves performance, and that annotation dimension specific prompt design improves performance. We release tweet IDs and labels along with our code at https://github.com/KathPra/Codebooks2VLMs.git.

视觉语言模型气候传播社交媒体分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。