arXiv:2606.16209cs.DLcs.IR2026-06

通过对比图文预训练模型,从150万张历史照片中识别出跨报纸传播的重印图像。

Viral Images: Identifying Reprintings within 1.5 Million Photographs in Chronicling America

论文配图:Viral Images: Identifying Reprintings within 1.5 Million Photographs in Chronicling America
图 1 · 摘自论文原文
  • 利用CLIP模型将150万张照片嵌入向量空间,无监督聚类发现重印内容。
  • 在1600万页报纸中识别出跨时间、跨报纸广泛传播的图像与广告集群。
  • 提供交互式网页界面,供人文研究者分析历史视觉文化的传播规律。

在'美国纪事'(Chronicling America)项目数百万数字化的美国历史报纸中,包含数千万张照片、插图、漫画和广告。这些视觉内容在不同报纸之间频繁重复出现。正如文本的重印反映了信息的传播,视觉内容的重印同样揭示了报纸作为信息流通与交换场所的特性。本文提出'病毒图像'(Viral Images)项目,旨在从150万张照片中识别重印内容。我们采用来自超过1600万页报纸的《报纸导航器》(Newspaper Navigator)数据集,引入一种基于对比语言-图像预训练(CLIP)的无监督方法,将照片嵌入向量空间并进行聚类,以发现重印图像。我们还开发了公开交互界面(https://viral-images.org),支持人文学者浏览与研究这些聚类结果。进一步分析揭示了跨越不同时期与报纸广泛传播的多样化图像与广告,为理解历史视觉文化的传播机制提供了新视角。

原文摘要 · Abstract (English)

Within the millions of digitized historic American newspapers in the Chronicling America initiative are tens of millions of photographs, illustrations, cartoons, and advertisements. Much of this visual culture is shared across newspaper titles and issues. Just as reprinted texts within these newspapers speak to the virality of textual content, so too does this reprinted visual culture speak to newspapers as sites of constant information circulation and exchange. In this paper, we introduce Viral Images, a project to identify reprintings within 1.5 million photographs in Chronicling America. For our analysis, we adopt the Newspaper Navigator dataset of extracted photographs from over 16 million pages in Chronicling America. We introduce an unsupervised method of identifying reprintings by leveraging contrastive language-image pretraining (CLIP) to embed these 1.5 million photographs and applying clustering to identify re-printed content. We detail our public interface, https://viral-images.org, which we designed in order to enable humanists to interactively browse and study these identified clusters. In addition, we analyze the identified clusters, uncovering a diversity of photographs and advertisements that have been circulated across different newspapers over time.

图像识别历史文献CLIP数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。