arXiv:2506.09847cs.CLcs.AI2025-06

构建带来源元数据的新闻数据集,用于检测图像与文本是否真实匹配。

Dataset of News Articles with Provenance Metadata for Media Relevance Assessment

  • 基于新闻文章和带来源标签的图像构建新数据集
  • 零样本下地点相关性任务表现良好,时间相关性任务仍有不足
  • 适合研究虚假信息检测、媒体可信度评估的研究者

如今,脱离语境和错误归属的图像已成为虚假信息与误导性传播的主要形式。现有检测方法通常仅关注图像语义是否与文本叙述相符,一旦图像内容与叙事存在部分对应,便难以发现伪造。为此,我们提出新闻媒体来源数据集(News Media Provenance Dataset),包含带有来源标记的新闻文章与图像。在此数据集上,我们定义了两个任务:起源地点相关性(LOR)与起源时间/日期相关性(DTOR),并对六种大型语言模型进行了基线测试。结果表明,尽管在零样本条件下LOR任务表现尚可,但DTOR任务性能受限,提示需设计专用架构并开展进一步研究。

原文摘要 · Abstract (English)

Out-of-context and misattributed imagery is the leading form of media manipulation in today's misinformation and disinformation landscape. The existing methods attempting to detect this practice often only consider whether the semantics of the imagery corresponds to the text narrative, missing manipulation so long as the depicted objects or scenes somewhat correspond to the narrative at hand. To tackle this, we introduce News Media Provenance Dataset, a dataset of news articles with provenance-tagged images. We formulate two tasks on this dataset, location of origin relevance (LOR) and date and time of origin relevance (DTOR), and present baseline results on six large language models (LLMs). We identify that, while the zero-shot performance on LOR is promising, the performance on DTOR hinders, leaving room for specialized architectures and future work.

虚假信息检测媒体可信度多模态分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。