arXiv:2410.06314cs.CVcs.CL2024-10

基于历史报纸数据集,提出时序图文检索新任务并发布竞赛结果

Temporal Image Caption Retrieval Competition -- Description and Results

  • 构建包含274年历史新闻的时序图文数据集
  • 竞赛验证了多模态模型在长时序文本-图像匹配中的性能
  • 适合对历史数据挖掘与跨模态检索感兴趣的学者

多模态模型融合视觉与文本信息,近年来受到广泛关注。本文针对文本-图像检索这一多模态挑战,引入一项新任务,将模态扩展至时间维度。本文介绍的时序图文检索竞赛(Temporal Image Caption Retrieval Competition, TICRC)基于Chronicling America和Challenging America项目,提供了一个涵盖274年数字化美国报纸的庞大历史数据集。除竞赛结果外,本文还分析了所用数据集的构成及其构建过程。

原文摘要 · Abstract (English)

Multimodal models, which combine visual and textual information, have recently gained significant recognition. This paper addresses the multimodal challenge of Text-Image retrieval and introduces a novel task that extends the modalities to include temporal data. The Temporal Image Caption Retrieval Competition (TICRC) presented in this paper is based on the Chronicling America and Challenging America projects, which offer access to an extensive collection of digitized historic American newspapers spanning 274 years. In addition to the competition results, we provide an analysis of the delivered dataset and the process of its creation.

时序检索历史数据多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。