arXiv:2505.14868cs.IRcs.CV2025-05被引 3

用视觉语义分析视频图像,自动提炼35类主题。

VisTopics: A Visual Semantic Unsupervised Approach to Topic Modeling of Video and Image Data

  • 结合LDA与字幕语义分析,端到端识别视觉主题。
  • 从1.1万帧视频中去重至6928帧,发现35个主题。
  • 适合媒体分析、传播研究者做跨平台叙事对比。

理解视觉叙事对分析媒体表征的演变至关重要。本文提出VisTopics,一种计算框架,通过端到端流程(帧提取、去重、语义聚类)分析大规模视觉数据。将该方法应用于452个NBC新闻视频,将11,070帧缩减至6,928帧去重后,通过语义分析揭示了35个主题,涵盖政治事件到环境危机等。结合隐含狄利克雷分布(LDA)与基于字幕的语义分析,VisTopics展示了在不同情境下揭示视觉框架模式的潜力。该方法支持纵向研究与跨平台比较,有助于理解媒体、技术与公共话语的交叉影响。通过人工编码准确率验证方法可靠性,并强调其在传播研究中的可扩展性。通过弥合视觉表征与语义意义之间的鸿沟,VisTopics为计算媒体研究提供了变革性工具。未来研究可利用它进行媒体机构或地理区域间的对比分析,揭示媒体叙事演变及其社会影响。

原文摘要 · Abstract (English)

Understanding visual narratives is crucial for examining the evolving dynamics of media representation. This study introduces VisTopics, a computational framework designed to analyze large-scale visual datasets through an end-to-end pipeline encompassing frame extraction, deduplication, and semantic clustering. Applying VisTopics to a dataset of 452 NBC News videos resulted in reducing 11,070 frames to 6,928 deduplicated frames, which were then semantically analyzed to uncover 35 topics ranging from political events to environmental crises. By integrating Latent Dirichlet Allocation with caption-based semantic analysis, VisTopics demonstrates its potential to unravel patterns in visual framing across diverse contexts. This approach enables longitudinal studies and cross-platform comparisons, shedding light on the intersection of media, technology, and public discourse. The study validates the method's reliability through human coding accuracy metrics and emphasizes its scalability for communication research. By bridging the gap between visual representation and semantic meaning, VisTopics provides a transformative tool for advancing the methodological toolkit in computational media studies. Future research may leverage VisTopics for comparative analyses across media outlets or geographic regions, offering insights into the shifting landscapes of media narratives and their societal implications.

主题建模视觉分析媒体研究无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。