用AI生成的假图配假文,骗过人类和模型,新方法能识破这种高阶造假。
The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
- 用MLLM生成与篡改图像语义一致的欺骗性文本,构建新型造假数据集。
- 在跨域测试中准确率达88.18%,mAP为60.25,可有效识别复杂伪造内容。
- 适合关注AI造假检测、多模态安全的研究者和应用开发者。
多媒体篡改检测已成为应对人工智能生成虚假信息的关键挑战。现有方法存在两大局限:(1) 低估了由多模态大语言模型(MLLM)驱动的欺骗风险——当前技术主要针对基于规则的文本篡改,却未能应对MLLM动态生成的、语义连贯且情境合理但具有误导性的叙事;(2) 现有场景依赖人为制造的不匹配内容,缺乏语义一致性,易被察觉。为此,我们提出一个全新对抗性流程,利用MLLM生成高风险虚假信息。首先构建MLLM驱动的合成多模态(MDSM)数据集:先使用先进编辑技术篡改图像,再配以与视觉篡改语义一致的MLLM生成文本。在此基础上,提出面向篡改诊断的MLLM框架AMD,包含两项创新:特征感知预编码策略和面向篡改的推理机制。全面实验验证了该框架在统一架构下检测MLLM驱动多模态欺骗的卓越泛化能力。在MDSM数据集的跨域测试中,平均准确率为88.18%,mAP为60.25,mIoU为61.02。
原文摘要 · Abstract (English)
The detection and grounding of multimedia manipulation has emerged as a critical challenge in combating AI-generated disinformation. While existing methods have made progress in recent years, we identify two fundamental limitations in current approaches: (1) Underestimation of MLLM-driven deception risk: prevailing techniques primarily address rule-based text manipulations, yet fail to account for sophisticated misinformation synthesized by multimodal large language models (MLLMs) that can dynamically generate semantically coherent, contextually plausible yet deceptive narratives conditioned on manipulated images; (2) Unrealistic misalignment artifacts: currently focused scenarios rely on artificially misaligned content that lacks semantic coherence, rendering them easily detectable. To address these gaps holistically, we propose a new adversarial pipeline that leverages MLLMs to generate high-risk disinformation. Our approach begins with constructing the MLLM-Driven Synthetic Multimodal (MDSM) dataset, where images are first altered using state-of-the-art editing techniques and then paired with MLLM-generated deceptive texts that maintain semantic consistency with the visual manipulations. Building upon this foundation, we present the Artifact-aware Manipulation Diagnosis via MLLM (AMD) framework featuring two key innovations: Artifact Pre-perception Encoding strategy and Manipulation-Oriented Reasoning, to tame MLLMs for the MDSM problem. Comprehensive experiments validate our framework's superior generalization capabilities as a unified architecture for detecting MLLM-powered multimodal deceptions. In cross-domain testing on the MDSM dataset, AMD achieves the best average performance, with 88.18 ACC, 60.25 mAP, and 61.02 mIoU scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。