构建首个真实世界多模态假信息定位数据集,支持跨模态误导内容识别。
A New Dataset and Benchmark for Grounding Multimodal Misinformation
- 提出跨模态假信息定位任务,融合文本、语音与视觉标注。
- 创建包含360个真实案例的GroundLie360数据集,经Snopes验证并附推理过程。
- 设计基于视觉语言模型的QA驱动检测方法,提升可解释性与定位精度。
在线虚假信息视频的泛滥带来严重社会风险。现有数据集和检测方法主要聚焦二分类或单模态定位,依赖后处理数据,缺乏解释能力以应对具有说服力的虚假内容。本文提出多模态假信息定位任务(GroundMM),旨在验证多模态内容并定位跨模态的误导片段。我们发布了首个真实世界数据集GroundLie360,涵盖虚假信息类型分类体系,对文本、语音和视觉进行细粒度标注,并通过Snopes证据和标注者推理进行验证。同时提出基于视觉语言模型的问答驱动基线方法FakeMark,利用单模态与跨模态线索实现高效检测与定位。实验揭示了该任务的挑战,为可解释的多模态假信息检测奠定基础。
原文摘要 · Abstract (English)
The proliferation of online misinformation videos poses serious societal risks. Current datasets and detection methods primarily target binary classification or single-modality localization based on post-processed data, lacking the interpretability needed to counter persuasive misinformation. In this paper, we introduce the task of Grounding Multimodal Misinformation (GroundMM), which verifies multimodal content and localizes misleading segments across modalities. We present the first real-world dataset for this task, GroundLie360, featuring a taxonomy of misinformation types, fine-grained annotations across text, speech, and visuals, and validation with Snopes evidence and annotator reasoning. We also propose a VLM-based, QA-driven baseline, FakeMark, using single- and cross-modal cues for effective detection and grounding. Our experiments highlight the challenges of this task and lay a foundation for explainable multimodal misinformation detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。