扩充数据集,专攻图像背景与主体不匹配的全局伪造问题
DGM4+: Dataset Extension for Global Scene Inconsistency
- 用提示词生成真人置于荒诞背景的新闻图,模拟真实伪造场景
- 新增5000样本,涵盖背景错位、文字篡改等三类全局伪造模式
- 适合研究多模态检测模型鲁棒性或反虚假信息的学者使用
生成模型的快速发展降低了制造逼真多模态虚假信息的门槛。伪造图像与篡改标题频繁结合,形成具有说服力的虚假叙事。尽管DGM4数据集为该领域研究奠定了基础,但其仅覆盖局部篡改(如人脸替换、属性修改、标题变更)。当前存在关键空白:全局不一致(如主体与背景不符)在真实伪造中日益普遍。为此,我们扩展DGM4,构建包含5000个高质量样本的DGM4+数据集,引入前景-背景(FG-BG)不匹配及其与文本篡改的混合形式。利用OpenAI的gpt-image-1和精心设计的提示词,生成以真人为主角、置于荒谬或不可能背景中的新闻风格图像(例如:教师在火星表面平静授课)。标题在三种条件下生成:字面、文本属性、文本拆分,分别对应三种新篡改类型:FG-BG、FG-BG+TA、FG-BG+TS。质量控制流程包括每图1至3张可见人脸、感知哈希去重、基于OCR的文字清理以及符合真实新闻标题长度的约束。通过引入全局篡改,本扩展补充现有数据集,构建可测试局部与全局推理能力的基准DGM4+。该资源旨在提升对多模态模型(如HAMMER)在处理FG-BG不一致时性能的评估。数据集与生成脚本已开源于https://github.com/Gaganx0/DGM4plus。
原文摘要 · Abstract (English)
The rapid advances in generative models have significantly lowered the barrier to producing convincing multimodal disinformation. Fabricated images and manipulated captions increasingly co-occur to create persuasive false narratives. While the Detecting and Grounding Multi-Modal Media Manipulation (DGM4) dataset established a foundation for research in this area, it is restricted to local manipulations such as face swaps, attribute edits, and caption changes. This leaves a critical gap: global inconsistencies, such as mismatched foregrounds and backgrounds, which are now prevalent in real-world forgeries. To address this, we extend DGM4 with 5,000 high-quality samples that introduce Foreground-Background (FG-BG) mismatches and their hybrids with text manipulations. Using OpenAI's gpt-image-1 and carefully designed prompts, we generate human-centric news-style images where authentic figures are placed into absurd or impossible backdrops (e.g., a teacher calmly addressing students on the surface of Mars). Captions are produced under three conditions: literal, text attribute, and text split, yielding three new manipulation categories: FG-BG, FG-BG+TA, and FG-BG+TS. Quality control pipelines enforce one-to-three visible faces, perceptual hash deduplication, OCR-based text scrubbing, and realistic headline length. By introducing global manipulations, our extension complements existing datasets, creating a benchmark DGM4+ that tests detectors on both local and global reasoning. This resource is intended to strengthen evaluation of multimodal models such as HAMMER, which currently struggle with FG-BG inconsistencies. We release our DGM4+ dataset and generation script at https://github.com/Gaganx0/DGM4plus
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。