用预训练图像嵌入构建可解释的政治图像主题模型
An Image is Worth $K$ Topics: A Visual Structural Topic Model with Pretrained Image Embeddings
- 融合预训练图像嵌入与结构化主题模型,捕捉政治图像语义
- 能识别出可解释且与网络政治传播相关的主题
- 适合研究社交媒体视觉内容的政治分析者使用
政治科学家越来越关注大规模视觉内容分析。然而,现有计算工具仍难以应对社会政治研究的特定挑战与目标。本文提出一种视觉结构化主题模型(vSTM),将预训练图像嵌入与结构化主题模型相结合。该方法具有显著优势:首先,预训练嵌入使模型能够捕捉与政治语境相关图像的语义复杂性;其次,结构化主题模型可分析主题与协变量之间的关系,同时保持图像作为多主题混合的精细表征。在实证应用中,vSTM成功识别出具有可解释性、一致性和实质性意义的主题,适用于在线政治传播研究。
原文摘要 · Abstract (English)
Political scientists are increasingly interested in analyzing visual content at scale. However, the existing computational toolbox is still in need of methods and models attuned to the specific challenges and goals of social and political inquiry. In this article, we introduce a visual Structural Topic Model (vSTM) that combines pretrained image embeddings with a structural topic model. This has important advantages compared to existing approaches. First, pretrained embeddings allow the model to capture the semantic complexity of images relevant to political contexts. Second, the structural topic model provides the ability to analyze how topics and covariates are related, while maintaining a nuanced representation of images as a mixture of multiple topics. In our empirical application, we show that the vSTM is able to identify topics that are interpretable, coherent, and substantively relevant to the study of online political communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。