用大模型分析图文视频元数据,识别新闻中内容是否被歪曲使用。
Large Language Models and Provenance Metadata for Determining the Relevance of Images and Videos in News Stories
- 结合文本与媒体元数据,判断图文视频在新闻中的相关性
- 能发现脱离上下文或伪造的多模态信息
- 适合做假新闻检测的科研人员和媒体审核者
最有效的虚假信息传播是多模态的,常将脱离上下文或完全虚构的图像、视频与文字结合以支持特定叙事。当前的假信息检测方法,无论是针对深度伪造还是文本文章,往往忽略多种模态之间的相互作用。本文提出的系统基于大语言模型,分析文章文本及所含图像、视频的出处元数据,判断其相关性。我们开源了系统原型及交互式网页界面。
原文摘要 · Abstract (English)
The most effective misinformation campaigns are multimodal, often combining text with images and videos taken out of context -- or fabricating them entirely -- to support a given narrative. Contemporary methods for detecting misinformation, whether in deepfakes or text articles, often miss the interplay between multiple modalities. Built around a large language model, the system proposed in this paper addresses these challenges. It analyzes both the article's text and the provenance metadata of included images and videos to determine whether they are relevant. We open-source the system prototype and interactive web interface.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。