arXiv:2505.15489cs.CVcs.CL2025-05被引 8

构建欺骗意图数据集,提升模型识别新闻误导性叙事能力

Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models

  • 用意图引导框架生成1.2万组图文对,模拟创作者误导策略
  • 14个主流视觉语言模型在意图推理上表现差,依赖表面线索
  • 数据集可检测误导意图、溯源来源、推断创作者动机

多模态虚假信息的影响不仅源于事实错误,更来自创作者刻意嵌入的误导性叙事。理解此类创作意图对多模态虚假信息检测(MMD)与信息治理至关重要。为此,我们提出DeceptionDecoded,一个包含12,000个图像-标题对的大规模基准数据集,基于可信参考文章构建,采用意图引导模拟框架,刻画新闻创作者的预期影响与执行策略。该数据集涵盖视觉与文本模态的多种操纵手段,支持三项以意图为中心的任务:(1) 误导意图检测,(2) 误导源归属,(3) 创作者意图推断。我们评估了14个先进视觉语言模型(VLMs),发现其在意图推理上表现不佳,常依赖表面一致性、风格修饰或启发式真实感信号。我们的框架通过系统合成数据,使模型学会深层意图推理。在DeceptionDecoded上训练的模型在真实世界MMD任务中表现出强泛化能力,验证了该框架既是诊断VLM脆弱性的基准,也是提升现实多模态虚假信息治理鲁棒性的高质量数据生成引擎。

原文摘要 · Abstract (English)

The impact of multimodal misinformation arises not only from factual inaccuracies but also from the misleading narratives that creators deliberately embed. Interpreting such creator intent is therefore essential for multimodal misinformation detection (MMD) and effective information governance. To this end, we introduce DeceptionDecoded, a large-scale benchmark of 12,000 image-caption pairs grounded in trustworthy reference articles, created using an intent-guided simulation framework that models both the desired influence and the execution plan of news creators. The dataset captures both misleading and non-misleading cases, spanning manipulations across visual and textual modalities, and supports three intent-centric tasks: (1) misleading intent detection, (2) misleading source attribution, and (3) creator desire inference. We evaluate 14 state-of-the-art vision-language models (VLMs) and find that they struggle with intent reasoning, often relying on shallow cues such as surface-level alignment, stylistic polish, or heuristic authenticity signals. To bridge this, our framework systematically synthesizes data that enables models to learn implication-level intent reasoning. Models trained on DeceptionDecoded demonstrate strong transferability to real-world MMD, validating our framework as both a benchmark to diagnose VLM fragility and a data synthesis engine that provides high-quality, intent-focused resources for enhancing robustness in real-world multimodal misinformation governance.

多模态虚假信息意图识别视觉语言模型数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。